Promptify: Text-to-Image Generation through Interactive Prompt Exploration with Large Language Models

Generative AI (Text, Image, Music, Video)Human-LLM CollaborationAI-Assisted Creative WritingUI/UX Designers

Title of the Paper

Promptify: Text-to-Image Generation through Interactive Prompt Exploration with Large Language Models

Paper Information

  • Domain: Text-to-image generation and prompt engineering in human-computer interaction
  • Keywords: Text-to-image generation, prompt engineering, large language models, interactive design, AI image generation

Research Background and Problem Statement

  • Identified Problems or Challenges:

    1. While text-to-image generation models (e.g., Stable Diffusion and DALL-E 2) are highly powerful, users face difficulties in crafting effective prompts to express complex creative ideas.
    2. Prompt crafting often involves a tedious trial-and-error process, and the systems lack built-in features to help users discover relevant keywords.
    3. Novice users are unfamiliar with related keywords and prompt design, making it challenging to effectively manage large volumes of randomly generated images.
  • Significance: Prompt engineering is a critical task in this domain, as it directly impacts the quality and style of generated images. Addressing the challenges in prompt crafting can significantly enhance the efficiency of text-to-image generation models, especially in creative fields.

  • Motivation and Related Work:

    1. Existing research on prompt engineering provides only high-level suggestions and does not cater to specific aesthetic goals of different users. Additionally, tools like Reddit communities and third-party resources share partial information but fail to meet personalized needs.
    2. Through user interviews, this paper further identifies challenges in prompt crafting, randomness management, and analyzing large batches of generated images.

Solution

  • Method or Solution: The authors propose Promptify, an interactive system for prompt exploration and optimization that integrates large language models (LLMs) with the open-source Stable Diffusion model. Promptify offers the following features:

    • Automatic prompt suggestions: Expands and generates complex prompts based on user-input themes and styles.
    • Image layout and classification: Groups and visualizes images by similarity to help users organize and browse generated images.
    • Image-based optimization suggestions: Extracts specific keywords from generated images to assist users in refining initial prompts.
  • Innovative Aspects:

    1. Utilizes GPT-3.5 as the prompt suggestion engine to separately expand themes and styles, allowing users to dynamically adjust prompt generation directions using natural language.
    2. Combines CLIP Embedding for image layout and clustering, enabling users to efficiently identify visual trends.
    3. For the first time, integrates LLM prompt techniques with text-to-image generation models to design a cross-modal prompt engineering tool.
  • Implementation Steps and Key Technologies:

    1. Employ zero-shot and few-shot prompt techniques to guide LLMs in generating complex prompt keywords.
    2. Use CLIP and TSNE mechanisms for visual layout and clustering of images, enabling labeling and user interaction.
    3. Extract keywords from generated images using CLIP Interrogator to guide prompt optimization.

Research Outcomes

  • Specific Results:

    1. Significantly reduced the cognitive load for users in crafting initial prompts, enabling novice users to achieve higher-quality image generation on their first attempt.
    2. Automatically extracted keywords to support users in iterative prompt optimization, enhancing visual presentation goals.
    3. Developed a complete interactive user interface to efficiently manage and compare large batches of generated images.
  • Advantages:

    1. Compared to existing tools (e.g., Automatic1111), Promptify significantly improves the efficiency and user experience of prompt engineering.
    2. Users reported lower mental demands and frustration during image generation.
  • Experimental or Evaluation Results:

    • User studies indicate that Promptify outperforms baseline tools in terms of aesthetic quality, consistency of generated images, and reduced cognitive load.
    • Most users found Promptify's theme and style expansion, image layout, clustering features, and keyword suggestions highly practical and praised their utility.
  • Limitations and Future Directions:

    1. The provided keyword suggestions rely on existing artist databases, which may pose usability challenges for users unfamiliar with relevant artistic references.
    2. The current system does not support negative prompt generation; future work could explore automatic generation of negative keywords to further optimize image outputs.
    3. Randomness management in prompt generation still requires technical improvement, such as investigating the impact of fragmented prompts on generation details.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/uist/126717/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3586183.3606725
At a Glance

Paper Snapshot

fact_check
dataset
Source
UIST
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Human-LLM Collaboration, AI-Assisted Creative Writing
work
Professions
UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers