SketchFlex: Facilitating Spatial-Semantic Coherence in Text-to-Image Generation with Region-Based Sketches

Honorable Mention
Generative AI (Text, Image, Music, Video)Human-LLM CollaborationSoftware Engineers & DevelopersUI/UX DesignersFreelancers (Design, Writing, Translation)

Research Background and Problem

  • Identified Problems or Challenges:
    Existing Text-to-Image Generation Models (T2I) can produce high-quality images, but non-expert users still face difficulties in using rough sketches to control the semantic and spatial coherence of generated images. These challenges are mainly reflected in:

    1. Users find it difficult to create appropriate text prompts based on sketches, especially for describing semantic relationships in multi-object scenes.
    2. The quality of rough sketches is generally low, leading to generated images that lack semantic logic, such as missing objects, unnatural relationships, and incorrect scenes.
    3. Most existing models are end-to-end, lacking iterative adjustment mechanisms, making it impossible for users to optimize specific regions individually.
  • Significance:
    Using rough sketches as control inputs provides non-expert users with an easy-to-operate generation method. However, issues with accuracy and coherence limit its practical application. Addressing this pain point is critical for improving the usability and efficiency of interactive generation tools, especially in creative design and artistic generation fields.

  • Research Motivation and Related Work:
    The primary motivation is to meet the needs of non-expert users for generating complex and semantically coherent images using sketches and simple prompts. Existing research mostly focuses on optimizing text prompts or spatial condition control but often requires high-precision inputs or fails to meet the needs of iterative generation, lacking support for rough user inputs.

Solution

  • Proposed Method/Solution:
    The authors propose an interactive system called SketchFlex, which allows users to generate high-quality and semantically coherent images using rough region sketches and simple text prompts. The system includes the following core modules:

    1. Sketch-Aware Prompt Suggestions: A multimodal large language model (MLLM) is used to recommend initial prompts that match the sketch, reducing the user's cognitive load.
    2. Spatial Condition Sketch Optimization: Rough user-drawn sketches are refined into high-quality spatial conditions to ensure the accuracy of the generated results.
    3. Iterative Interaction Interface: A user interface is designed to support sketch control and prompt editing, enabling users to iteratively improve images across multiple generations.
  • Innovations:

    1. Integration of user sketches and text prompts to automatically generate semantically comprehensive prompts.
    2. Proposal of a "Decompose-and-Recompose Strategy" to iteratively optimize individual object sketches and integrate spatial conditions for multiple objects.
    3. Provision of a user-friendly interactive interface for non-expert users, supporting sketch region adjustments and multiple generation iterations.
  • Implementation Steps and Key Techniques:

    1. Create a semantic space defining key semantic dimensions required for image generation, including object types, attributes, states, and spatial and semantic relationships among multiple objects.
    2. Extract user input sketches and initial prompts, and use multimodal models and external datasets (e.g., Visual Genome and VAW) to generate suggested prompts.
    3. Apply the Decompose-and-Recompose Strategy to refine individual object sketches, generate candidate shapes, and combine user adjustments into final image anchor conditions.
    4. Utilize Canny edge detectors and the ControlNet model to achieve precise shape anchoring, ensuring consistency between generated images and user intent.

Research Outcomes

  • Specific Outcomes:

    1. SketchFlex can generate semantically coherent and high-quality images, significantly reducing users' cognitive load.
    2. A user study demonstrated that SketchFlex aligns better with user intent and provides high controllability and flexibility compared to existing methods.
    3. Experimental results show that the system outperforms existing end-to-end generation models and region-based generation models in multi-object image generation.
  • Advantages Over Existing Solutions:

    1. Users can iteratively generate and modify specific image objects without needing to regenerate the entire image.
    2. The system supports free sketch adjustments and prompt refinement, significantly improving the generation quality for non-expert users.
    3. Optimized prompt suggestions lead to more stable and semantically accurate generation results.
  • Experimental or Evaluation Results:

    1. User satisfaction surveys indicate that SketchFlex significantly outperforms baseline models (T2I and R2I) in terms of image quality, intent alignment, and image coherence.
    2. In two tasks, the system demonstrated higher accuracy and consistency, producing results closest to reference images.
  • Limitations and Future Directions:

    1. When dealing with three or more semantically similar objects, object omission may occur.
    2. The system cannot precisely interpret fine and thin sketch strokes, limiting its usability for professional users.
    3. The model's generation capability still needs improvement for more complex spatial semantic relationships (e.g., containment relationships).
    4. The authors suggest future directions, including supporting user-uploaded custom models to enhance customization, designing more sophisticated filtering mechanisms to avoid inappropriate content, and seamlessly integrating sketch refinement with generation interactions.

In summary, SketchFlex demonstrates excellent performance in enhancing non-expert users' image generation capabilities and provides valuable insights for designing interactive multimodal generation systems.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188525/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713801
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
Honorable Mention
group
Authors
4 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Human-LLM Collaboration
work
Professions
Software Engineers & Developers, UI/UX Designers, Freelancers (Design, Writing, Translation)
article
Content Status
Full text indexed
hub
Related Papers
10 related papers