Examining the Text-to-Image Community of Practice: Why and How do People Prompt Generative AIs?
Honorable MentionAuthors
Image generation entered the mainstream with the spread of large machine learning (ML) models capable of generating artistic images from text. AI art enthusiasts generated millions of images online within a few months, revealing the future social, economic, and legal consequences of text-to-image (TTI) generation. This work explores the multifaceted sociology and practices behind this yet understudied phenomenon. We aim to understand text-to-image practitioners, their motivations, practice, and the usability challenges they face in order to envision creative and meaningful interactions with this technology. We analyzed two sets of data, an online questionnaire answered by 64 practitioners gathered on social media groups and DiffusionDB, a large dataset of user text prompts sent to the Stable Diffusion generative model. The questionnaire results suggest that TTI generation is a recreational activity from narrow socio-professional groups whose users employ auxiliary techniques across platforms and beyond request-response interactions. Inherent model limitations and finding suitable prompt formulation are their main problems. The analysis of the DiffusionDB dataset informed the creation of a taxonomy for prompt specifiers and a corresponding model capable of recognizing the semantic content of unseen prompts, which further informed how practitioners structure TTI prompt. We finally discuss the design and socio-technical implications of our research for creativity support.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
Sketchforme: Composing Sketched Scenes from Text Descriptions for Interactive Applications
UIST '19· Generative AI (Text, Image, Music, Video) +2
- 83%
StyleFactory: Towards Better Style Alignment in Image Creation through Style-Strength-Based Control and Evaluation
UIST '24· Generative AI (Text, Image, Music, Video) +2
- 80%
GANravel: User-Driven Direction Disentanglement in Generative Adversarial Networks
CHI '23· Generative AI (Text, Image, Music, Video) +1
- 80%
Artinter: AI-powered Boundary Objects for Commissioning Visual Arts
DIS '23· Generative AI (Text, Image, Music, Video) +1
- 67%
"I don't want to feel like I'm working in a 1960s factory": The Practitioner Perspective on Creativity Support Tool Adoption
CHI '22· Generative AI (Text, Image, Music, Video) +1
- 67%
TypeDance: Creating Semantic Typographic Logos from Image through Personalized Generation
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 67%
When Teams Embrace AI: Human Collaboration Strategies in Generative Prompting in a Creative Design Task
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 67%
InkIdeator: Supporting Chinese-Style Visual Design Ideation via AI-Infused Exploration of Chinese Paintings
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 67%
Evolving Roles and Workflows of Creative Practitioners in the Age of Generative AI
C&C '24· Generative AI (Text, Image, Music, Video) +1
- 67%
GANCollage: A GAN-Driven Digital Mood Board to Facilitate Ideation in Creativity Support
DIS '23· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)