FusAIn: Composing Generative AI Visual Prompts Using Pen-based Interaction

Generative AI (Text, Image, Music, Video)UI/UX DesignersVisual Artists & Designers

Research Background and Problem

  • Identified Problems or Challenges: Despite the powerful capabilities of current generative AI (GenAI) in image generation, it primarily relies on text input and full-image interactions, limiting designers' creative engagement with visual materials. Expressing subtle design intentions and creating professional design assets using existing GenAI tools remain challenging.
  • Significance: These issues hinder the broader application of GenAI in professional design practices, making it difficult for designers to fully control the generated outcomes. Designers require fine-grained control and editability to ensure alignment between content and design intentions.
  • Research Motivation and Related Work: Previous studies have emphasized the necessity of providing designers with more controllable and expressive interactions. However, current tools are fragmented and lack unified process support. Additionally, while extensive work has explored engineering and optimization in text-to-image generation, there is a lack of focused exploration on visual-based input for generation.

Solution

  • Proposed Method/Solution:
    • Introduced a generative AI visual prompting tool FusAIn, based on the concept of an "intelligent" pen.
    • Enables designers to extract visual attributes such as objects, colors, and textures through pen interactions and compose user-defined visual prompts based on these attributes.
    • Provides both local and global generation modes, with support for re-editing generated content.
  • Innovations:
    • Developed a new interaction mechanism, the "intelligent pen," which integrates design intentions with AI generation.
    • Broke the dominance of text as input in generative AI, proposing a new pen-based visual "composition as prompt" interaction paradigm.
    • Emphasized the deconstruction and reconstruction of design materials, making the generated results more controllable and editable.
  • Implementation Steps and Techniques:
    1. Tool Design and Technical Implementation:
      • Developed the FusAIn tool's user interface and pen functionality modules, including object pen, color pen, and texture pen.
      • Integrated the Stable Diffusion generation model and ControlNet technology to support fine-grained control of visual attributes.
    2. Interaction Support:
      • Enabled the pen to extract and apply different attributes, such as color picking, texture overlay, and object manipulation.
      • Allowed switching between local (selected area-based) and global (entire canvas-based) generation modes.
    3. User Feedback:
      • Built interface components such as canvas, loading tools, editing panels, and style locking to ensure designers can iterate on designs throughout the process.

Research Outcomes

  • Specific Outcomes:
    • FusAIn improved designers' ability to define visual details at different levels, providing fine-grained control to enhance the editability and reusability of generated images.
    • The introduced "intelligent pen" allowed users to easily manipulate complex visual details, aligning closely with designers' workflows.
  • Advantages over Existing Solutions:
    • Compared to traditional AI tools relying on text prompts (e.g., Adobe Firefly), FusAIn achieved higher alignment with intentions and controllability through its canvas-based display and comprehensive pen interaction features.
    • The visual prompt interaction significantly reduced the likelihood of errors (e.g., misusing layers) while fostering designers' creativity.
  • Experimental or Evaluation Results:
    • User Study: Conducted comparative experiments with 12 professional designers, demonstrating that FusAIn significantly enhanced the ability to express complex visual details and match design intentions.
      • Over 75% of users preferred using FusAIn for visual creation.
      • Participants found pen interactions to provide a highly precise and intuitive way of visual expression, free from language constraints.
    • Qualitative Feedback: Designers felt greater control over the generation process and considered the system more suitable for creativity-driven tasks, such as transitioning from sketches to final designs.
  • Limitations and Future Directions:
    • Limitations: The quality of generated image details needs improvement; limited user study time may not have covered all functionalities; fixed model parameters may not perform optimally across all tasks.
    • Future Directions:
      • Explore expanding the application of the "intelligent pen" in complex fields such as architecture and fashion.
      • Introduce more complex visual attribute interactions (e.g., environment or ambiance) to broaden the design scope.
      • Optimize the "intelligent pen" application to meet the needs of users with varying levels of experience.

Through this research, FusAIn reveals the potential of visual-based generative AI inputs, providing designers with an intuitive, precise, and creativity-supportive interaction form. This study offers significant insights for the future design of generative AI tools.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188881/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714027
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video)
work
Professions
UI/UX Designers, Visual Artists & Designers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers