Is It AI or Is It Me? Understanding Users’ Prompt Journey with Text-to-Image Generative AI Tools
Authors
Title of the Paper
Is It AI or Is It Me? Understanding Users’ Prompt Journey with Text-to-Image Generative AI Tools
Paper Information
- Research Area: Human-Computer Interaction, Generative AI, User Experience Research
- Keywords: Prompt Engineering, Generative AI, Text-to-Image Generation, User Journey, Human-Computer Interaction, Creativity Tools, Midjourney, AI Prompt Optimization
Research Background and Issues
-
Identified Problems or Challenges:
- Users of text-to-image generation tools (e.g., Midjourney) often struggle to achieve their design goals directly when crafting prompts. The complexity and diversity of prompts remain major obstacles in user adoption.
- Existing research on prompt guidelines, manuals, and "prompt engineering" tends to be overly theoretical, failing to fully address users' real-world goals and practical challenges.
- There is a mismatch between generative AI outputs and user goals, requiring users to overcome a steep learning curve to effectively utilize these tools.
-
Significance and Motivation:
- As generative AI becomes increasingly prevalent in creative fields (e.g., graphic design, architectural design), understanding how users intuitively design prompts is crucial for improving tool design.
- Examining specific cases of how users create, evaluate, and refine prompts can help better align tool capabilities with users' creative goals, while also advancing user education and community engagement for generative AI tools.
-
Research Questions:
- What is the typical prompt design process for users of text-to-image generation tools?
- What are the main obstacles users face during the prompt design process?
Solutions
-
Methods or Solutions:
- Conducted semi-structured interviews to study the prompt design processes and challenges of 19 Midjourney users.
- Proposed a "User Prompt Journey Model," which includes specific strategies for the three stages: prompt structure design, image evaluation, and prompt refinement.
-
Innovations:
- Identified five common prompt structures: descriptive sentences, templates, overview + details, segmented patterns, and word sequences.
- Explained how users evaluate and optimize generated images based on the goals and content types of each stage.
- Discovered that prompt creation is both an individual and a collaborative social activity, emphasizing the roles of "collaborative prompt creation" and "prompt learning" within the generative AI community.
-
Implementation Steps and Techniques:
- The interview process was divided into two phases: user background surveys and playback analysis of prompt and generation cases.
- Data analysis employed thematic analysis to synthesize interview content and extract patterns and challenges.
- Used the Midjourney image generation tool to analyze its input-output process and users' strategies for prompt optimization.
Research Outcomes
-
Specific Findings:
- Prompt Structures: Characteristics and applicable scenarios for five prompt structure strategies (e.g., "descriptive sentences" are suitable for beginners, while "segmented patterns" offer greater control over details).
- Evaluation Criteria: Standards for evaluating generated images are divided into two dimensions: content and visual design, including metrics such as realism, detail, color, and composition.
- Optimization Strategies: Users employ various prompt optimization methods, such as adding descriptive terms, removing unnecessary words, changing word order, adjusting weight parameters, and regenerating outputs.
-
Comparison with Existing Solutions:
- This research integrates users' prompt journeys with community learning, viewing interaction behaviors as both an exploration of individual experiences and an extension of social learning. This perspective is novel in the literature.
- The study goes beyond theoretical modeling by deeply exploring users' practical operations and psychological processes.
-
Experimental or Evaluation Results:
- The prompt journey is a dynamic, iterative process where users continuously adjust their goals during creation.
- Users with different levels of experience face distinct challenges: beginners rely more on random generation, while experienced users prefer to control prompt structures.
-
Limitations and Future Directions:
- Limitations:
- This study focuses solely on the Midjourney tool and does not cover differences in prompt patterns across other generative AI tools.
- The interview participants were from a single linguistic and cultural background, limiting the study's applicability to cross-cultural contexts.
- Future Directions:
- Conduct comparative studies on prompt patterns across different generative AI tools.
- Expand the scope of research to include more diverse and cross-cultural user backgrounds.
- Explore how community-driven "prompt learning" systems influence user creativity and tool improvement.
- Limitations:
Conclusion
Through detailed user interview analysis, this study reveals the strategies and challenges faced by Midjourney users in prompt structure design, image evaluation, and optimization processes, providing critical guidance for human-centered design of generative AI tools. This research lays a theoretical foundation for the future development of more intelligent, personalized, and user-friendly text-to-image generation systems.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
2- What is the typical prompt construction workflow for text-to-image tool users?Category: Generative Image Creation and Editing ControlSimilar questionsarrow_forward
- What are the main obstacles users face during prompt construction?Category: Generative Image Creation and Editing ControlSimilar questionsarrow_forward
Practical Problems
1- Users often struggle to achieve desired effects when crafting prompts with generative AI tools.Category: Generative Image Creation and Editing ControlSimilar questionsarrow_forward
- 80%
IntentTuner: An Interactive Framework for Integrating Human Intentions in Fine-tuning Text-to-Image Generative Models
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 80%
AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose Tools
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 67%
MUD: Towards a Large-Scale and Noise-Filtered UI Dataset for Modern Style UI Modeling
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 67%
How the Role of Generative AI Shapes Perceptions of Value in Human-AI Collaborative Work
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 67%
GANzilla: User-Driven Direction Discovery in Generative Adversarial Networks
UIST '22· Generative AI (Text, Image, Music, Video) +1
- 67%
Knoll: Creating a Knowledge Ecosystem for Large Language Models
UIST '25· Human-LLM Collaboration +1
- 60%
User Experience Design Professionals’ Perceptions of Generative Artificial Intelligence
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 60%
Cells, Generators, and Lenses: Design Framework for Object-Oriented Interaction with Large Language Models
UIST '23· Human-LLM Collaboration
- 60%
Patchview: LLM-powered Worldbuilding with Generative Dust and Magnet Visualization
UIST '24· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)