Brickify: Enabling Expressive Design Intent Specification through Direct Manipulation on Design Tokens

Generative AI (Text, Image, Music, Video)Graphic Design & Typography ToolsCreative Collaboration & Feedback SystemsUI/UX DesignersProduct Designers

Research Background and Problem Statement

  • Identified Issues and Challenges:

    1. Current methods of expressing design intent through natural language prompts (e.g., text) are challenging. Designers must condense and accurately convey ambiguous visual details, a process that is both cumbersome and unintuitive.
    2. While AI tools (e.g., DALL·E, MidJourney) excel in generating aesthetically pleasing images, they lack the element control required by designers, particularly in complex design tasks where relationships between elements and iterative refinements need to be explicitly expressed.
    3. Existing tools rely heavily on textual descriptions to construct complex visual representations. This approach is limited by the discrete nature of language in expressing continuous attributes (e.g., size, spatial positioning) and requires significant effort for repeated inputs (e.g., rephrasing reference images).
  • Research Significance:

    1. Designers not only demand high-quality outputs but also require control over elements during the design decision-making process, especially in professional design scenarios where relationships between multiple elements must be explicitly defined.
    2. Expressing design language visually (rather than textually) may align more closely with designers' mental models, reducing communication overhead.
    3. Facilitating efficient collaboration between AI and humans enables AI to better understand designers' intentions.
  • Research Motivation and Related Work:

    1. The limitations of current tools stem from their linear, top-down approach, which fails to effectively support the resolution of complex visual queries.
    2. Related fields have developed "step-by-step text prompting" (modular prompting), cross-modal semantic understanding tools, and direct visual manipulation tools. However, these approaches remain constrained in addressing "how elements are constructed into a whole."
    3. Thus, a new visual-driven interaction paradigm is needed to enrich and enhance the flexibility of visual and interactive expression of design intent.

Solution

  • Methodology: The authors propose a visual-driven interaction paradigm called Brickify, which emphasizes "direct manipulation on design tokens" to help users clearly express required design elements and their relationships.

    • Design Tokens: Visual elements (e.g., subjects, colors, styles) are concretized into interactive graphical representations.
    • Construction and Manipulation of Design Intent: Users intuitively express relationships and spatial structures between visual tokens by dragging, repositioning, resizing, grouping, and connecting tokens.
    • Intent Execution Mechanism: Advanced AI models such as BoxDiff and Style-Align translate the user's visual lexicon into control signals, enabling the generation of final images.
  • Innovations:

    1. Introduced a visual-driven interaction paradigm that allows designers to directly manipulate segments within reference images rather than the entire image.
    2. Defined three key token types: visual tokens (specific elements), text tokens (supplementary textual descriptions), and imagination tokens (providing AI with greater creative freedom).
    3. Introduced persistent and temporary token management to support efficient reuse and iterative design of elements, a feature difficult to achieve with current text-prompt interactions.
    4. Improved the interpretability of visual representations, enhanced human-AI collaboration efficiency, and expanded the flexibility of element construction and reuse.
  • Implementation Steps and Key Technologies:

    1. Token Generation:
      • Use techniques (e.g., SAM model and Break-A-Scene method) to segment and extract specific subjects, colors, or styles from reference images, converting them into tokens.
      • Provide automated and manual color extraction features, style transfer tools, etc.
    2. Token Manipulation:
      • Support dragging, moving, resizing, grouping, and connecting tokens, allowing cross-referencing between tokens.
      • Enable iterative adjustments of token visual structures, helping users explore different design options.
    3. Intent Execution:
      • Use BoxDiff to constrain layout control in AI image generation.
      • Employ Style-Align to synchronize visual styles, HistoGAN for global color adjustment, and Blended Latent Diffusion for local color adjustment.
    4. Provide a history panel and step-by-step feedback mechanism to track interaction results and support localized debugging.

Research Outcomes

  • Key Results:

    1. Improved User Experience: Experiments demonstrated that designers could more efficiently express complex design intentions using Brickify, especially in tasks involving numerous elements and intricate relationships.
    2. User Feedback (Study 1): Participants reported significantly better experiences in terms of "clarity of expression" and "reduced cognitive load" compared to traditional text prompts.
    3. Creativity Support Index (CSI): Brickify scored highly in supporting exploratory design and providing a pleasant user experience.
  • Experimental Evaluation:

    1. Task Time Analysis: While the initial setup with Brickify took slightly longer, the refinement phase was significantly shorter, particularly in complex design tasks (Hard condition).
    2. External Evaluation Results: Compared to text prompts, Brickify demonstrated significantly higher clarity in element precision and the representation of spatial and color relationships.
    3. User Preferences: 88% of users expressed a preference for using Brickify in the future.
  • Limitations:

    1. Brickify performs less effectively in tasks requiring entirely novel creations (non-combinatorial designs) and highly freeform expressions.
    2. The high computational cost of backend technologies (e.g., AI model inference) limits the possibility of real-time feedback, potentially affecting creative flow.
    3. The study included only experienced designers, so the results may not fully generalize to novice designers.
  • Future Directions:

    1. Expand token types, such as camera angles and textures, to broaden the range of expressions.
    2. Support user-defined token types and reconfiguration features to address diverse design needs.
    3. Develop a bidirectional editable visual lexicon extraction system, enabling both users and AI to collaboratively construct editable design elements.
    4. Explore Brickify's potential in dynamic video production or 3D scene modeling beyond static design.

Conclusion

The introduction of Brickify redefines the core question of "how to express design intent more intuitively and efficiently" in human-computer collaboration, granting AI more controllable input execution capabilities. This work holds significant value in the field of interface interaction and provides robust theoretical and practical support for future explorations in designing visual creativity tools.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188912/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714087
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Graphic Design & Typography Tools, Creative Collaboration & Feedback Systems
work
Professions
UI/UX Designers, Product Designers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers