Preference-Guided Prompt Optimization for Text-to-Image Generation
Paper Title
Preference-Guided Prompt Optimization for Text-to-Image Generation
Publication Info
- Topic area: Human-centered prompt optimization for generative models.
- Keywords: Text-to-image generation, prompt optimization, user preferences, generative models, exploration-exploitation balance, adaptive optimization, binary feedback, human-AI collaboration.
Background and Problem
- Problem / challenge: Crafting effective prompts for generative models is challenging due to the implicit nature of user goals, the opaque generative process, and the cognitive demand of manual refinement. Existing methods either rely heavily on human effort or numerical reward functions, which are unsuitable for human-centered tasks requiring sparse feedback and rapid convergence.
- Significance: Reducing user effort in prompt optimization while achieving satisfactory generation outcomes is critical for making generative models accessible and effective in creative tasks.
- Motivation and related work: Prior approaches include manual prompt refinement tools, automatic optimization using numerical functions, and iterative refinement incorporating user feedback. However, these methods impose high cognitive load or fail to generalize to diverse user goals. A gap remains in integrating automatic optimization with minimal user input, such as binary preference feedback.
Solution
- Proposed approach: Adaptive Preferential Prompt Optimization (APPO), a method that refines prompts based solely on user preference feedback to guide text-to-image generation tasks.
- Novelty:
- APPO introduces three complementary strategies: retainment, alignment, and expansion, to balance exploration and exploitation.
- An adaptive policy dynamically adjusts exploration intensity based on semantic similarity between prompts.
- APPO minimizes user effort by relying on binary preference feedback rather than manual edits or numerical ratings.
- APPO ensures essential information from initial prompts is preserved through a consistency check mechanism.
- Procedure and key techniques:
- Retainment: Preserves preferred prompts from previous iterations to maintain stability.
- Alignment: Uses textual gradients to refine non-preferred prompts toward user preferences.
- Expansion: Applies evolutionary algorithms (crossover and mutation) to explore diverse directions in the prompt space.
- Adaptive policy: Adjusts mutation intensity based on semantic similarity to balance exploration and exploitation.
- Consistency check: Ensures essential elements from the initial prompt are retained in optimized prompts.
Results
- Concrete findings:
- APPO achieved a 19.64% improvement in synthetic tests compared to 9.38–14.66% for ablated variants.
- Outperformed TextGrad (6.45%) and Self-TICK (8.06%) in synthetic tests.
- Reduced iterations and time required for user satisfaction in user studies (average <4 iterations for APPO vs. >6 for baselines).
- Advantage over baselines:
- APPO demonstrated faster convergence, higher output quality, and lower cognitive load compared to PromptCharm, DSPy, and Clarification methods.
- Participants reported lower mental and physical demand, less frustration, and higher satisfaction with APPO.
- Experiments / evaluation:
- Synthetic tests: Compared APPO and variants under controlled scenarios using simulated user feedback.
- User study: 16 participants completed close- and open-ended image generation tasks across four conditions (APPO, PromptCharm, DSPy, Clarification). Metrics included iterations, time, NASA-TLX workload, and CSI scores.
- Limitations and future work:
- Struggles with highly specific or overly complex user intents.
- Limited applicability to tasks requiring evolving goals or missing essential elements.
- Future directions include integrating richer user feedback, extending to multi-modal inputs, and leveraging shared preferences for meta-learning.
Summary
APPO is a human-centered prompt optimization method that refines prompts based on binary preference feedback, enabling users to achieve satisfactory text-to-image generation outcomes with minimal effort. By balancing exploration and exploitation through retainment, alignment, and expansion strategies, and employing an adaptive policy, APPO effectively converges on high-quality prompts while preserving essential information. Synthetic tests and user studies demonstrate its superiority over existing methods in terms of efficiency, user satisfaction, and cognitive load reduction. APPO represents a step toward seamless human-AI collaboration in creative generative tasks, with potential extensions to other domains and multi-modal inputs.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
Partnering with Generative AI: Experimental Evaluation of Model-Led and Human-Led Interaction in Human-AI Co-Creation
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 83%
I Lead, You Help But Only with Enough Details: Understanding User Experience of Co-Creation with Artificial Intelligence
CHI '18· Generative AI (Text, Image, Music, Video) +2
- 83%
``Control Is a Trajectory, Not a Point'': Conceptualizing Control in Human-AI Co-Creativity
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 83%
Studying Collaborative Interactive Machine Teaching in Image Classification
IUI '24· Human-LLM Collaboration +1
- 83%
GANzilla: User-Driven Direction Discovery in Generative Adversarial Networks
UIST '22· Generative AI (Text, Image, Music, Video) +1
- 71%
AI-Augmented Brainwriting: Investigating the use of LLMs in group ideation
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 71%
Design Principles for Generative AI Applications
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 71%
Understanding Nonlinear Collaboration between Human and AI Agents: A Co-design Framework for Creative Design
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 71%
Collaborative Document Editing with Multiple Users and AI Agents
CHI '26· Human-LLM Collaboration +2
- 71%
Less Redraw, More Explore: Suggestion and Completion for Sketch-to-Image
CHI '26· Generative AI (Text, Image, Music, Video) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)