Brickify: Enabling Expressive Design Intent Specification through Direct Manipulation on Design Tokens
Authors
Research Background and Problem Statement
-
Identified Issues and Challenges:
- Current methods of expressing design intent through natural language prompts (e.g., text) are challenging. Designers must condense and accurately convey ambiguous visual details, a process that is both cumbersome and unintuitive.
- While AI tools (e.g., DALL·E, MidJourney) excel in generating aesthetically pleasing images, they lack the element control required by designers, particularly in complex design tasks where relationships between elements and iterative refinements need to be explicitly expressed.
- Existing tools rely heavily on textual descriptions to construct complex visual representations. This approach is limited by the discrete nature of language in expressing continuous attributes (e.g., size, spatial positioning) and requires significant effort for repeated inputs (e.g., rephrasing reference images).
-
Research Significance:
- Designers not only demand high-quality outputs but also require control over elements during the design decision-making process, especially in professional design scenarios where relationships between multiple elements must be explicitly defined.
- Expressing design language visually (rather than textually) may align more closely with designers' mental models, reducing communication overhead.
- Facilitating efficient collaboration between AI and humans enables AI to better understand designers' intentions.
-
Research Motivation and Related Work:
- The limitations of current tools stem from their linear, top-down approach, which fails to effectively support the resolution of complex visual queries.
- Related fields have developed "step-by-step text prompting" (modular prompting), cross-modal semantic understanding tools, and direct visual manipulation tools. However, these approaches remain constrained in addressing "how elements are constructed into a whole."
- Thus, a new visual-driven interaction paradigm is needed to enrich and enhance the flexibility of visual and interactive expression of design intent.
Solution
-
Methodology: The authors propose a visual-driven interaction paradigm called Brickify, which emphasizes "direct manipulation on design tokens" to help users clearly express required design elements and their relationships.
- Design Tokens: Visual elements (e.g., subjects, colors, styles) are concretized into interactive graphical representations.
- Construction and Manipulation of Design Intent: Users intuitively express relationships and spatial structures between visual tokens by dragging, repositioning, resizing, grouping, and connecting tokens.
- Intent Execution Mechanism: Advanced AI models such as BoxDiff and Style-Align translate the user's visual lexicon into control signals, enabling the generation of final images.
-
Innovations:
- Introduced a visual-driven interaction paradigm that allows designers to directly manipulate segments within reference images rather than the entire image.
- Defined three key token types: visual tokens (specific elements), text tokens (supplementary textual descriptions), and imagination tokens (providing AI with greater creative freedom).
- Introduced persistent and temporary token management to support efficient reuse and iterative design of elements, a feature difficult to achieve with current text-prompt interactions.
- Improved the interpretability of visual representations, enhanced human-AI collaboration efficiency, and expanded the flexibility of element construction and reuse.
-
Implementation Steps and Key Technologies:
- Token Generation:
- Use techniques (e.g., SAM model and Break-A-Scene method) to segment and extract specific subjects, colors, or styles from reference images, converting them into tokens.
- Provide automated and manual color extraction features, style transfer tools, etc.
- Token Manipulation:
- Support dragging, moving, resizing, grouping, and connecting tokens, allowing cross-referencing between tokens.
- Enable iterative adjustments of token visual structures, helping users explore different design options.
- Intent Execution:
- Use BoxDiff to constrain layout control in AI image generation.
- Employ Style-Align to synchronize visual styles, HistoGAN for global color adjustment, and Blended Latent Diffusion for local color adjustment.
- Provide a history panel and step-by-step feedback mechanism to track interaction results and support localized debugging.
- Token Generation:
Research Outcomes
-
Key Results:
- Improved User Experience: Experiments demonstrated that designers could more efficiently express complex design intentions using Brickify, especially in tasks involving numerous elements and intricate relationships.
- User Feedback (Study 1): Participants reported significantly better experiences in terms of "clarity of expression" and "reduced cognitive load" compared to traditional text prompts.
- Creativity Support Index (CSI): Brickify scored highly in supporting exploratory design and providing a pleasant user experience.
-
Experimental Evaluation:
- Task Time Analysis: While the initial setup with Brickify took slightly longer, the refinement phase was significantly shorter, particularly in complex design tasks (Hard condition).
- External Evaluation Results: Compared to text prompts, Brickify demonstrated significantly higher clarity in element precision and the representation of spatial and color relationships.
- User Preferences: 88% of users expressed a preference for using Brickify in the future.
-
Limitations:
- Brickify performs less effectively in tasks requiring entirely novel creations (non-combinatorial designs) and highly freeform expressions.
- The high computational cost of backend technologies (e.g., AI model inference) limits the possibility of real-time feedback, potentially affecting creative flow.
- The study included only experienced designers, so the results may not fully generalize to novice designers.
-
Future Directions:
- Expand token types, such as camera angles and textures, to broaden the range of expressions.
- Support user-defined token types and reconfiguration features to address diverse design needs.
- Develop a bidirectional editable visual lexicon extraction system, enabling both users and AI to collaboratively construct editable design elements.
- Explore Brickify's potential in dynamic video production or 3D scene modeling beyond static design.
Conclusion
The introduction of Brickify redefines the core question of "how to express design intent more intuitively and efficiently" in human-computer collaboration, granting AI more controllable input execution capabilities. This work holds significant value in the field of interface interaction and provides robust theoretical and practical support for future explorations in designing visual creativity tools.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can a vision-driven interaction paradigm express design intent more intuitively and efficiently?Category: AI/LLM as Design Collaborators and Creative ToolsSimilar questionsarrow_forward
- Can visual token systems express element relationships and spatial structure more explicitly than text prompts?Category: AI/LLM as Design Collaborators and Creative ToolsSimilar questionsarrow_forward
- Can AI better understand designers' intent and generate high-quality output through user-manipulated visual tokens?Category: AI/LLM as Design Collaborators and Creative ToolsSimilar questionsarrow_forward
Practical Problems
1- Designers struggle to express detailed design intent through text prompts, which is difficult and time-consuming.Category: AI/LLM as Design Collaborators and Creative ToolsSimilar questionsarrow_forward
- 100%
CreativeConnect: Supporting Reference Recombination for Graphic Design Ideation with Generative AI
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 83%
FashionQ: An AI-Driven Creativity Support Tool for Facilitating Ideation in Fashion Design
CHI '21· Generative AI (Text, Image, Music, Video) +2
- 83%
Fashioning Creative Expertise with Generative AI: Graphical Interfaces for Design Space Exploration Better Support Ideation Than Text Prompts
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 83%
Exploring Interactive Color Palettes for Abstraction-Driven Exploratory Image Colorization
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 83%
Creativity from Surprise: Bridging the Gap Between Fashion Designers' Inspiration Work and AI Creative Support Tools
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 83%
DesignTrace: Exploring, Iterating and Tracking Design Alternatives with GenAI
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 83%
DesignPrompt: Using Multimodal Interaction for Design Exploration with Generative AI
DIS '24· Generative AI (Text, Image, Music, Video) +2
- 83%
Enhancing Generative AI Image Refinement with Scribbles and Annotations: A Comparative Study of Multimodal Prompts
IUI '26· Generative AI (Text, Image, Music, Video) +2
- 80%
VisiBlends: A Flexible Workflow for Visual Blends
CHI '19· Graphic Design & Typography Tools +1
- 80%
Learning Personal Style from Few Examples
DIS '21· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)