RePrompt: Automatic Prompt Editing to Refine AI-Generative Art Towards Precise Expressions
Authors
Title of the Paper
RePrompt: Automatic Prompt Editing to Refine AI-Generative Art Towards Precise Expressions
Paper Information
- Research Area: Human-Computer Interaction (HCI), Text-to-Image Generation Models, Emotional Expression
- Keywords: Text-to-Image Generation, Prompt Engineering, AI-Generated Visual Art, Emotional Expression, Explainable AI
Research Background and Problem
-
Identified Problems or Challenges:
- Existing text-to-image generation models, such as DALL·E 2, can generate images based on free-text input, but the generated images often fail to accurately reflect the emotions or context expressed in the input text.
- Users typically need to iteratively modify text prompts through trial and error to improve image outputs. This process is inefficient and lacks systematic support.
- The effectiveness of AI generative models in expressing complex emotions has not been fully validated.
-
Why This Problem is Important:
- Emotional expression is a vital aspect of human social interaction and mental health. Supporting users in quickly generating emotionally expressive content has significant application potential.
- Enhancing the emotional accuracy of AI-generated images can help users better express themselves and create through visual art.
-
Research Motivation and Related Work:
- The scientific community has begun to focus on prompt engineering for text-to-image generation models, but existing work primarily relies on a limited set of fixed keywords, lacking complex and diverse natural expressions.
- Subjective evaluation methods for emotional expression in images have limitations, and the impact of emotions on visual art generation remains an underexplored area.
Solution
-
Proposed Method or Solution: The authors developed an automated prompt editing method called RePrompt, designed to refine text prompts to optimize the emotional accuracy of generated images.
-
Innovative Aspects:
- Designed based on users' intuitive strategies for prompt editing, combined with explainable AI techniques to train a proxy model.
- Provides a clear set of text editing rules (Rubric) to enable intuitive understanding and operation for both users and the model.
- Applies explainable AI techniques to prompt engineering, helping AI generate better content while improving human control and understanding of AI processes.
-
Implementation Steps and Key Techniques:
- Feature Selection: Extract parts of speech (e.g., nouns, verbs, adjectives) from the text and intuitive features related to context, such as word concreteness.
- Proxy Model Training: Use VQGAN-CLIP to generate 10,000 images and train a lightweight proxy model by calculating image emotion alignment scores (CLIP Score).
- Model Interpretation: Use SHAP to analyze the feature contributions of the proxy model and generate optimized feature value ranges.
- Rubric Creation for Prompt Editing: Based on feature analysis results, establish rules for automated editing, such as adding or removing words to optimize the number or concreteness of nouns.
- Automated Prompt Editing: Apply the rules to generate optimized text prompts and use text-to-image generation models to create images.
Research Outcomes
-
Specific Outcomes:
- Effectiveness Validation:
- The prompt editing method significantly improved the emotional accuracy of generated images (particularly for negative emotions).
- Achieved better alignment between images and text.
- User Evaluation:
- User evaluation experiments showed that prompts generated by RePrompt better facilitated emotional expression, especially in scenarios involving negative emotions.
- Effectiveness Validation:
-
Comparative Advantages Over Existing Solutions:
- Compared to simply adding emotion tags, RePrompt significantly enhances the emotional accuracy of image generation.
- Does not rely on complex and hard-to-understand features or techniques, offering better explainability and user acceptability.
-
Experimental or Evaluation Results:
- By comparing the quality of generated images under different conditions through computational and user studies, RePrompt-generated images outperformed those created with original prompts and other prompt engineering methods.
- Found that two types of CLIP scores (text-to-image alignment and image-to-emotion alignment) were more significant in predicting negative emotions but weaker in supporting positive emotions.
-
Limitations and Future Directions:
- Limitations:
- Does not account for the overall sentence structure, potentially overlooking the contextual meaning of specific phrases.
- The data and model have weaker recognition capabilities for positive emotions, requiring more precise training datasets.
- The model used (CLIP) may have biases in modeling positive emotions.
- Future Directions:
- Combine other emotion detection tools or physical signals, such as EEG, to more comprehensively measure emotional expression in images.
- Develop stronger modeling and dataset expansion for positive emotions.
- Apply the RePrompt framework to other generative models (e.g., music or video generation) to explore prompt editing systems for different types of expressions.
- Limitations:
Conclusion
RePrompt leverages explainable AI and prompt engineering to provide automated emotional optimization for text-to-image generation models, achieving significant advancements in negative emotion expression and model transparency. It also opens up broad opportunities for further development in human-AI collaboration and generative AI applications in mental health.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can AI-generated images more accurately reflect the emotions and context in input text?Category: Generative Image Creation and Editing ControlSimilar questionsarrow_forward
- How can automated text prompt editing improve emotional accuracy and consistency in image generation?Category: Generative Image Creation and Editing ControlSimilar questionsarrow_forward
- How can negative emotions be effectively expressed in AI-generated visual art?Category: Generative Image Creation and Editing ControlSimilar questionsarrow_forward
Practical Problems
1- Users struggle to precisely express complex emotions visually through existing AI generation tools.Category: Generative Image Creation and Editing ControlSimilar questionsarrow_forward
- 83%
Varif.ai to Vary and Verify User-Driven Diversity in Scalable Image Generation
DIS '25· Generative AI (Text, Image, Music, Video) +2
- 80%
FlatMagic: Improving Flat Colorization through AI-driven Design for Digital Comic Professionals
CHI '22· Generative AI (Text, Image, Music, Video) +1
- 80%
When is a Tool a Tool? User Perceptions of System Agency in Human-AI Co-Creative Drawing
DIS '23· Generative AI (Text, Image, Music, Video) +1
- 67%
The Effects of Generative AI on Design Fixation and Divergent Thinking
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 60%
Design Guidelines for Prompt Engineering Text-to-Image Generative Models
CHI '22· Generative AI (Text, Image, Music, Video)
- 60%
FusAIn: Composing Generative AI Visual Prompts Using Pen-based Interaction
CHI '25· Generative AI (Text, Image, Music, Video)
- 60%
SwipeGANSpace: Swipe-to-Compare Image Generation via Efficient Latent Space Exploration
IUI '24· Generative AI (Text, Image, Music, Video)
Based on Jaccard similarity of research subtopics & professions (≥60%)