PromptCharm: Text-to-Image Generation through Multi-modal Prompting and Refinement
Authors
Generative AI (Text, Image, Music, Video)Explainable AI (XAI)AI-Assisted Creative WritingSoftware Engineers & DevelopersUI/UX DesignersVisual Artists & Designers
Title of the Paper
PromptCharm: Text-to-Image Generation through Multi-modal Prompting and Refinement
Paper Information
- Research Area: Human-Computer Interaction, Generative AI, Text-to-Image Generation
- Keywords: Generative AI, Prompt Engineering, Large Language Models, Text-to-Image Generation, Human-AI Interaction, Model Attention
Research Background and Problem Statement
- What problems or challenges did the authors identify?
- Current generative AI models (e.g., Stable Diffusion) perform well in "text-to-image generation," but for novice users, effectively crafting and optimizing prompts to generate high-quality images that meet expectations remains a significant challenge.
- Existing systems lack user-friendly feedback mechanisms, such as explaining the correspondence between generated images and prompts or effectively bridging the gap between model outputs and user intent.
- Why is this problem important?
- Prompt engineering is critical to the performance of generative models, and optimizing the generation process is an essential step in improving the practical application of these models.
- As generative AI becomes more widespread, lowering the barrier to crafting prompts and enhancing the interactive experience can enable more non-expert users to engage with this technology.
- Research Motivation and Related Work:
- Current research primarily focuses on automated prompt optimization (e.g., the Promptist model) and exploring interactive designs for prompt engineering (e.g., PromptMagician, PromptPaint). There is a need to further integrate model interpretability and multimodal interaction to enhance user experience.
Solution
- What methods or solutions did the authors propose?
- The authors proposed and designed a hybrid interactive system—PromptCharm—to assist users in optimizing prompts, generating images, and iterating feedback through multimodal interactions.
- The main features of PromptCharm include automatic prompt optimization, image style exploration, model attention visualization and adjustment, and image inpainting.
- What are the innovative aspects of this solution?
- The system integrates multimodal interaction and provides a rich feedback loop to address the potential gap between generated content and user intent.
- By visualizing and adjusting model attention, the system allows users to directly manipulate generated images rather than relying on cumbersome prompt modifications.
- What are the implementation steps and key technologies used?
- Prompt Optimization: Utilized Microsoft's Promptist model to automatically optimize the initial prompts provided by users.
- Image Style Exploration: Leveraged a frequency-counting method on tagged data from the DiffusionDB dataset to extract popular "modifiers" for style selection.
- Model Attention Visualization and Adjustment: Used the DAAM model interpretation technique to map generated content to the textual components of the prompt.
- Image Inpainting: Combined Stable Diffusion's inpainting functionality, allowing users to mark areas for replacement and generate modifications directly.
Research Outcomes
- What specific outcomes were achieved?
- PromptCharm significantly reduced the learning curve for users in crafting and optimizing prompts, enabling non-expert users to generate high-quality images that align with their expectations.
- Users experienced greater control and gained a more intuitive understanding of the generated results through model attention maps.
- What advantages does it have compared to existing solutions?
- PromptCharm outperformed baseline tools and existing automated optimization tools (e.g., Promptist) by enhancing interaction and feedback, helping users achieve significantly better generation results.
- With attention adjustment and image inpainting, users could address discrepancies between generated content and their intent without frequently rewriting prompts.
- What were the experimental or evaluation results?
- Two user studies were conducted:
- Closed-task study: Participants achieved an average structural similarity index (SSIM) of 0.648 between generated images and target images using PromptCharm, significantly higher than the baseline tool (0.479) and Promptist (0.574).
- Open-task study: Participants rated PromptCharm-generated images as more aesthetically pleasing and better aligned with their expectations, with a median Likert score of 6.0 out of 7.
- Two user studies were conducted:
- Limitations and Future Directions:
- Limitations: When generating large-area image inpainting, the patched regions may not blend seamlessly with the overall image. Advanced features like attention control are challenging to implement for closed-source generative models (e.g., DALL-E, MidJourney).
- Future Directions: Improve inpainting algorithm performance, support more generative models, expand to larger-scale image generation, and explore multilingual adaptation and detailed explanations for text modifiers.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How can the learning curve for non-expert users generating high-quality text-to-image results be reduced?Category: Generative Image Creation and Editing ControlSimilar questionsarrow_forward
- How can multimodal interaction improve users' understanding of the relationship between generated results and text prompts in GenAI?Category: Generative Image Creation and Editing ControlSimilar questionsarrow_forward
- Can model attention visualization and adjustment improve user satisfaction with text-to-image generation results?Category: Generative Image Creation and Editing ControlSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Non-expert users struggle to write and optimize prompts to achieve ideal text-to-image generation.Category: Generative Image Creation and Editing ControlSimilar questionsarrow_forward
- 71%
Prompting for Discovery: Flexible Sense-Making for AI Art-Making with Dreamsheets
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 71%
Understanding User Perceptions and the Role of AI Image Generators in Image Creation Workflows
CHI '25· Generative AI (Text, Image, Music, Video) +2
- 67%
FlatMagic: Improving Flat Colorization through AI-driven Design for Digital Comic Professionals
CHI '22· Generative AI (Text, Image, Music, Video) +1
- 67%
When is a Tool a Tool? User Perceptions of System Agency in Human-AI Co-Creative Drawing
DIS '23· Generative AI (Text, Image, Music, Video) +1
- 63%
POET: Supporting Prompting Creativity and Personalization with Automated Expansion of Text-to-Image Generation
UIST '25· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642803
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Explainable AI (XAI), AI-Assisted Creative Writing
work
Professions
Software Engineers & Developers, UI/UX Designers, Visual Artists & Designers
article
Content Status
Full text indexed
hub
Related Papers
5 related papers