Promptify: Text-to-Image Generation through Interactive Prompt Exploration with Large Language Models
Authors
Title of the Paper
Promptify: Text-to-Image Generation through Interactive Prompt Exploration with Large Language Models
Paper Information
- Domain: Text-to-image generation and prompt engineering in human-computer interaction
- Keywords: Text-to-image generation, prompt engineering, large language models, interactive design, AI image generation
Research Background and Problem Statement
-
Identified Problems or Challenges:
- While text-to-image generation models (e.g., Stable Diffusion and DALL-E 2) are highly powerful, users face difficulties in crafting effective prompts to express complex creative ideas.
- Prompt crafting often involves a tedious trial-and-error process, and the systems lack built-in features to help users discover relevant keywords.
- Novice users are unfamiliar with related keywords and prompt design, making it challenging to effectively manage large volumes of randomly generated images.
-
Significance: Prompt engineering is a critical task in this domain, as it directly impacts the quality and style of generated images. Addressing the challenges in prompt crafting can significantly enhance the efficiency of text-to-image generation models, especially in creative fields.
-
Motivation and Related Work:
- Existing research on prompt engineering provides only high-level suggestions and does not cater to specific aesthetic goals of different users. Additionally, tools like Reddit communities and third-party resources share partial information but fail to meet personalized needs.
- Through user interviews, this paper further identifies challenges in prompt crafting, randomness management, and analyzing large batches of generated images.
Solution
-
Method or Solution: The authors propose Promptify, an interactive system for prompt exploration and optimization that integrates large language models (LLMs) with the open-source Stable Diffusion model. Promptify offers the following features:
- Automatic prompt suggestions: Expands and generates complex prompts based on user-input themes and styles.
- Image layout and classification: Groups and visualizes images by similarity to help users organize and browse generated images.
- Image-based optimization suggestions: Extracts specific keywords from generated images to assist users in refining initial prompts.
-
Innovative Aspects:
- Utilizes GPT-3.5 as the prompt suggestion engine to separately expand themes and styles, allowing users to dynamically adjust prompt generation directions using natural language.
- Combines CLIP Embedding for image layout and clustering, enabling users to efficiently identify visual trends.
- For the first time, integrates LLM prompt techniques with text-to-image generation models to design a cross-modal prompt engineering tool.
-
Implementation Steps and Key Technologies:
- Employ zero-shot and few-shot prompt techniques to guide LLMs in generating complex prompt keywords.
- Use CLIP and TSNE mechanisms for visual layout and clustering of images, enabling labeling and user interaction.
- Extract keywords from generated images using CLIP Interrogator to guide prompt optimization.
Research Outcomes
-
Specific Results:
- Significantly reduced the cognitive load for users in crafting initial prompts, enabling novice users to achieve higher-quality image generation on their first attempt.
- Automatically extracted keywords to support users in iterative prompt optimization, enhancing visual presentation goals.
- Developed a complete interactive user interface to efficiently manage and compare large batches of generated images.
-
Advantages:
- Compared to existing tools (e.g., Automatic1111), Promptify significantly improves the efficiency and user experience of prompt engineering.
- Users reported lower mental demands and frustration during image generation.
-
Experimental or Evaluation Results:
- User studies indicate that Promptify outperforms baseline tools in terms of aesthetic quality, consistency of generated images, and reduced cognitive load.
- Most users found Promptify's theme and style expansion, image layout, clustering features, and keyword suggestions highly practical and praised their utility.
-
Limitations and Future Directions:
- The provided keyword suggestions rely on existing artist databases, which may pose usability challenges for users unfamiliar with relevant artistic references.
- The current system does not support negative prompt generation; future work could explore automatic generation of negative keywords to further optimize image outputs.
- Randomness management in prompt generation still requires technical improvement, such as investigating the impact of fragmented prompts on generation details.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- What frustrations hinder users when designing effective prompts for text-to-image generation?Category: Generative Image Creation and Editing ControlSimilar questionsarrow_forward
- How can an interactive system integrating large language models (e.g., GPT-3.5) with Stable Diffusion optimize users' prompt creation?Category: Generative Image Creation and Editing ControlSimilar questionsarrow_forward
- Can combining text and image cross-modal techniques improve users' experience in managing and reviewing large batches of images?Category: Generative Image Creation and Editing ControlSimilar questionsarrow_forward
Practical Problems
1- Users struggle to design prompts suitable for generating complex creative images.Category: Generative Image Creation and Editing ControlSimilar questionsarrow_forward
- 80%
PhraseFlow: Designs and Empirical Studies of Phrase-Level Input
CHI '21· Generative AI (Text, Image, Music, Video) +2
- 75%
User Experience Design Professionals’ Perceptions of Generative Artificial Intelligence
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 75%
Thoughtful, Confused, or Untrustworthy: How Text Presentation Influences Perceptions of AI Writing Tools
C&C '25· Generative AI (Text, Image, Music, Video) +2
- 75%
Patchview: LLM-powered Worldbuilding with Generative Dust and Magnet Visualization
UIST '24· Generative AI (Text, Image, Music, Video) +1
- 67%
ExpressEdit: Video Editing with Natural Language and Sketching
IUI '24· Generative AI (Text, Image, Music, Video) +3
- 60%
Sketching NLP: A Case Study of Exploring the Right Things To Design with Language Intelligence
CHI '19· Human-LLM Collaboration +1
- 60%
FlatMagic: Improving Flat Colorization through AI-driven Design for Digital Comic Professionals
CHI '22· Generative AI (Text, Image, Music, Video) +1
- 60%
Exploring Challenges and Opportunities to Support Designers in Learning to Co-create with AI-based Manufacturing Design Tools
CHI '23· Generative AI (Text, Image, Music, Video) +1
- 60%
OPTIMISM: Enabling Collaborative Implementation of Domain-Specific Metaheuristic Optimization
CHI '23· Generative AI (Text, Image, Music, Video) +1
- 60%
PopBlends: Strategies for Conceptual Blending with Large Language Models
CHI '23· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)