SketchFlex: Facilitating Spatial-Semantic Coherence in Text-to-Image Generation with Region-Based Sketches
Honorable MentionAuthors
Research Background and Problem
-
Identified Problems or Challenges:
Existing Text-to-Image Generation Models (T2I) can produce high-quality images, but non-expert users still face difficulties in using rough sketches to control the semantic and spatial coherence of generated images. These challenges are mainly reflected in:- Users find it difficult to create appropriate text prompts based on sketches, especially for describing semantic relationships in multi-object scenes.
- The quality of rough sketches is generally low, leading to generated images that lack semantic logic, such as missing objects, unnatural relationships, and incorrect scenes.
- Most existing models are end-to-end, lacking iterative adjustment mechanisms, making it impossible for users to optimize specific regions individually.
-
Significance:
Using rough sketches as control inputs provides non-expert users with an easy-to-operate generation method. However, issues with accuracy and coherence limit its practical application. Addressing this pain point is critical for improving the usability and efficiency of interactive generation tools, especially in creative design and artistic generation fields. -
Research Motivation and Related Work:
The primary motivation is to meet the needs of non-expert users for generating complex and semantically coherent images using sketches and simple prompts. Existing research mostly focuses on optimizing text prompts or spatial condition control but often requires high-precision inputs or fails to meet the needs of iterative generation, lacking support for rough user inputs.
Solution
-
Proposed Method/Solution:
The authors propose an interactive system called SketchFlex, which allows users to generate high-quality and semantically coherent images using rough region sketches and simple text prompts. The system includes the following core modules:- Sketch-Aware Prompt Suggestions: A multimodal large language model (MLLM) is used to recommend initial prompts that match the sketch, reducing the user's cognitive load.
- Spatial Condition Sketch Optimization: Rough user-drawn sketches are refined into high-quality spatial conditions to ensure the accuracy of the generated results.
- Iterative Interaction Interface: A user interface is designed to support sketch control and prompt editing, enabling users to iteratively improve images across multiple generations.
-
Innovations:
- Integration of user sketches and text prompts to automatically generate semantically comprehensive prompts.
- Proposal of a "Decompose-and-Recompose Strategy" to iteratively optimize individual object sketches and integrate spatial conditions for multiple objects.
- Provision of a user-friendly interactive interface for non-expert users, supporting sketch region adjustments and multiple generation iterations.
-
Implementation Steps and Key Techniques:
- Create a semantic space defining key semantic dimensions required for image generation, including object types, attributes, states, and spatial and semantic relationships among multiple objects.
- Extract user input sketches and initial prompts, and use multimodal models and external datasets (e.g., Visual Genome and VAW) to generate suggested prompts.
- Apply the Decompose-and-Recompose Strategy to refine individual object sketches, generate candidate shapes, and combine user adjustments into final image anchor conditions.
- Utilize Canny edge detectors and the ControlNet model to achieve precise shape anchoring, ensuring consistency between generated images and user intent.
Research Outcomes
-
Specific Outcomes:
- SketchFlex can generate semantically coherent and high-quality images, significantly reducing users' cognitive load.
- A user study demonstrated that SketchFlex aligns better with user intent and provides high controllability and flexibility compared to existing methods.
- Experimental results show that the system outperforms existing end-to-end generation models and region-based generation models in multi-object image generation.
-
Advantages Over Existing Solutions:
- Users can iteratively generate and modify specific image objects without needing to regenerate the entire image.
- The system supports free sketch adjustments and prompt refinement, significantly improving the generation quality for non-expert users.
- Optimized prompt suggestions lead to more stable and semantically accurate generation results.
-
Experimental or Evaluation Results:
- User satisfaction surveys indicate that SketchFlex significantly outperforms baseline models (T2I and R2I) in terms of image quality, intent alignment, and image coherence.
- In two tasks, the system demonstrated higher accuracy and consistency, producing results closest to reference images.
-
Limitations and Future Directions:
- When dealing with three or more semantically similar objects, object omission may occur.
- The system cannot precisely interpret fine and thin sketch strokes, limiting its usability for professional users.
- The model's generation capability still needs improvement for more complex spatial semantic relationships (e.g., containment relationships).
- The authors suggest future directions, including supporting user-uploaded custom models to enhance customization, designing more sophisticated filtering mechanisms to avoid inappropriate content, and seamlessly integrating sketch refinement with generation interactions.
In summary, SketchFlex demonstrates excellent performance in enhancing non-expert users' image generation capabilities and provides valuable insights for designing interactive multimodal generation systems.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can non-expert users generate semantically coherent, visually high-quality images from coarse regional sketches and simple text prompts?Category: Generative Image Creation and Editing ControlSimilar questionsarrow_forward
- How can coarse sketches be optimized into high-quality spatial conditions to ensure generation accuracy?Category: Generative Image Creation and Editing ControlSimilar questionsarrow_forward
- Can an interface supporting stepwise iterative drawing and text editing significantly improve image-generation flexibility and user satisfaction?Category: Generative Image Creation and Editing ControlSimilar questionsarrow_forward
Practical Problems
1- Non-expert users struggle to generate semantically accurate and visually coherent images from coarse sketches.Category: Generative Image Creation and Editing ControlSimilar questionsarrow_forward
- 80%
Generative and Malleable User Interfaces with Generative and Evolving Task-Driven Data Model
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 80%
The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 80%
GeneyMAP: Exploring the Potential of GenAI to Facilitate Mapping User Journeys for UX Design
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 67%
PhraseFlow: Designs and Empirical Studies of Phrase-Level Input
CHI '21· Generative AI (Text, Image, Music, Video) +2
- 67%
Stylette: Styling the Web with Natural Language
CHI '22· Immersion & Presence Research +2
- 67%
MUD: Towards a Large-Scale and Noise-Filtered UI Dataset for Modern Style UI Modeling
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 67%
Shaping Human-AI Collaboration: Varied Scaffolding Levels in Co-writing with Language Models
CHI '24· Human-LLM Collaboration +1
- 67%
How the Role of Generative AI Shapes Perceptions of Value in Human-AI Collaborative Work
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 67%
agentAR: Creating Augmented Reality Applications with Tool-Augmented LLM-based Autonomous Agents
UIST '25· AR Navigation & Context Awareness +2
- 60%
User Experience Design Professionals’ Perceptions of Generative Artificial Intelligence
CHI '24· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)