RePrompt: Automatic Prompt Editing to Refine AI-Generative Art Towards Precise Expressions

Generative AI (Text, Image, Music, Video)AI-Assisted Creative WritingUI/UX DesignersAI/ML Researchers & EngineersVisual Artists & Designers

Title of the Paper

RePrompt: Automatic Prompt Editing to Refine AI-Generative Art Towards Precise Expressions

Paper Information

  • Research Area: Human-Computer Interaction (HCI), Text-to-Image Generation Models, Emotional Expression
  • Keywords: Text-to-Image Generation, Prompt Engineering, AI-Generated Visual Art, Emotional Expression, Explainable AI

Research Background and Problem

  • Identified Problems or Challenges:

    1. Existing text-to-image generation models, such as DALL·E 2, can generate images based on free-text input, but the generated images often fail to accurately reflect the emotions or context expressed in the input text.
    2. Users typically need to iteratively modify text prompts through trial and error to improve image outputs. This process is inefficient and lacks systematic support.
    3. The effectiveness of AI generative models in expressing complex emotions has not been fully validated.
  • Why This Problem is Important:

    1. Emotional expression is a vital aspect of human social interaction and mental health. Supporting users in quickly generating emotionally expressive content has significant application potential.
    2. Enhancing the emotional accuracy of AI-generated images can help users better express themselves and create through visual art.
  • Research Motivation and Related Work:

    1. The scientific community has begun to focus on prompt engineering for text-to-image generation models, but existing work primarily relies on a limited set of fixed keywords, lacking complex and diverse natural expressions.
    2. Subjective evaluation methods for emotional expression in images have limitations, and the impact of emotions on visual art generation remains an underexplored area.

Solution

  • Proposed Method or Solution: The authors developed an automated prompt editing method called RePrompt, designed to refine text prompts to optimize the emotional accuracy of generated images.

  • Innovative Aspects:

    1. Designed based on users' intuitive strategies for prompt editing, combined with explainable AI techniques to train a proxy model.
    2. Provides a clear set of text editing rules (Rubric) to enable intuitive understanding and operation for both users and the model.
    3. Applies explainable AI techniques to prompt engineering, helping AI generate better content while improving human control and understanding of AI processes.
  • Implementation Steps and Key Techniques:

    1. Feature Selection: Extract parts of speech (e.g., nouns, verbs, adjectives) from the text and intuitive features related to context, such as word concreteness.
    2. Proxy Model Training: Use VQGAN-CLIP to generate 10,000 images and train a lightweight proxy model by calculating image emotion alignment scores (CLIP Score).
    3. Model Interpretation: Use SHAP to analyze the feature contributions of the proxy model and generate optimized feature value ranges.
    4. Rubric Creation for Prompt Editing: Based on feature analysis results, establish rules for automated editing, such as adding or removing words to optimize the number or concreteness of nouns.
    5. Automated Prompt Editing: Apply the rules to generate optimized text prompts and use text-to-image generation models to create images.

Research Outcomes

  • Specific Outcomes:

    1. Effectiveness Validation:
      • The prompt editing method significantly improved the emotional accuracy of generated images (particularly for negative emotions).
      • Achieved better alignment between images and text.
    2. User Evaluation:
      • User evaluation experiments showed that prompts generated by RePrompt better facilitated emotional expression, especially in scenarios involving negative emotions.
  • Comparative Advantages Over Existing Solutions:

    1. Compared to simply adding emotion tags, RePrompt significantly enhances the emotional accuracy of image generation.
    2. Does not rely on complex and hard-to-understand features or techniques, offering better explainability and user acceptability.
  • Experimental or Evaluation Results:

    1. By comparing the quality of generated images under different conditions through computational and user studies, RePrompt-generated images outperformed those created with original prompts and other prompt engineering methods.
    2. Found that two types of CLIP scores (text-to-image alignment and image-to-emotion alignment) were more significant in predicting negative emotions but weaker in supporting positive emotions.
  • Limitations and Future Directions:

    1. Limitations:
      • Does not account for the overall sentence structure, potentially overlooking the contextual meaning of specific phrases.
      • The data and model have weaker recognition capabilities for positive emotions, requiring more precise training datasets.
      • The model used (CLIP) may have biases in modeling positive emotions.
    2. Future Directions:
      • Combine other emotion detection tools or physical signals, such as EEG, to more comprehensively measure emotional expression in images.
      • Develop stronger modeling and dataset expansion for positive emotions.
      • Apply the RePrompt framework to other generative models (e.g., music or video generation) to explore prompt editing systems for different types of expressions.

Conclusion

RePrompt leverages explainable AI and prompt engineering to provide automated emotional optimization for text-to-image generation models, achieving significant advancements in negative emotion expression and model transparency. It also opens up broad opportunities for further development in human-AI collaboration and generative AI applications in mental health.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/96020/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581402
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), AI-Assisted Creative Writing
work
Professions
UI/UX Designers, AI/ML Researchers & Engineers, Visual Artists & Designers
article
Content Status
Full text indexed
hub
Related Papers
7 related papers