Jigsaw: Supporting Designers to Prototype Multimodal Applications by Chaining AI Foundation Models

Generative AI (Text, Image, Music, Video)Human-LLM CollaborationPrototyping & User TestingUI/UX DesignersProduct Designers

Document Title

Jigsaw: Supporting Designers to Prototype Multimodal Applications by Assembling AI Foundation Models

Document Information

  • Domain: Human-Computer Interaction, Design and AI Integration, Visual Programming
  • Keywords: Prototyping, Machine Learning, Foundation Models, Multimodal, Visual Programming Interface

Research Background and Problem

  • Challenges:

    1. Designers have limited understanding of the full capabilities of AI foundation models, making it difficult to fully leverage their potential.
    2. Writing "AI-friendly" prompts and adjusting parameters is cumbersome.
    3. Integrating cross-platform, multimodal models is highly challenging.
    4. The prototyping process is slow, hindering effective rapid iteration and experimentation.
  • Significance: With the rapid development of AI foundation models, designers are attempting to integrate these models into creative workflows but face numerous technical and workflow barriers. Addressing these issues will enhance designers' efficiency and promote broader adoption of AI technologies.

  • Motivation: The authors aim to design a tool (Jigsaw) that allows designers to easily explore, combine, and utilize AI foundation models, thereby overcoming the aforementioned challenges.

Solution

  • Proposed Approach:

    1. Jigsaw uses puzzle pieces as a metaphor for foundation models, with each piece representing a function or model.
    2. Includes a "Catalog Panel" to display all available models, providing model descriptions, usage examples, and search functionality.
    3. Employs color-coded blocks and a snapping mechanism to clarify model compatibility, simplifying the combination of multimodal models.
    4. Integrates an "Assembly Assistant" feature to recommend model combinations based on natural language descriptions.
    5. Offers real-time feedback on intermediate results, enabling users to quickly observe and adjust during construction and debugging.
  • Innovations:

    1. An intuitive, puzzle-piece-based visual interface significantly lowers the technical barrier for designers using AI.
    2. Integrated prompt generation and parameter explanation features help users easily create effective AI model inputs.
    3. Seamless integration of multimodal models.
    4. A novel block-based visual programming interface tailored for non-technical designers.
  • Implementation Steps:

    1. Created a model library supporting text, images, video, 3D, audio, and sketches, comprising 39 models in total.
    2. Designed specific input, output types, and default parameters for each model.
    3. Implemented LLM-based "glue pieces" to support multi-model chaining tasks.
    4. Integrated real-time input and output panels, supporting multimodal input and output formats including text, images, and audio.

Research Outcomes

  • Key Results:

    1. Developed the Jigsaw system, simplifying the process for designers to use foundation models.
    2. Enabled users to construct complex AI-based design workflows through a puzzle-piece approach.
    3. Facilitated model exploration, rapid prototyping, and visualization of design records.
  • Advantages Over Existing Solutions:

    1. A block-based interface designed for ease of use by non-technical designers.
    2. Effectively guides designers in selecting suitable models and achieving model combinations.
    3. Supports cross-platform collaboration and rapid prototyping of multimodal models.
  • Experimental Results:

    • Conducted user studies with ten designers, revealing that Jigsaw significantly helps designers:
      1. Explore new AI capabilities and gain deeper understanding of models through semantic search and puzzle-piece prompts.
      2. Rapidly construct and iterate AI creative workflows.
      3. Intuitively combine models across multiple tasks and modalities.
  • Limitations and Future Directions:

    1. Currently, only one AI model is supported per task; future plans include expanding to multiple model options.
    2. Enhancing the Assembly Assistant's ability to handle complex tasks and support interactive iterations with designers.
    3. Expanding support for real-time video and audio stream input/output.
    4. Exploring the integration of multimodal large language model (MLLM) technology into Jigsaw.
    5. Creating a library of design templates to facilitate experience and workflow sharing among designers.

Conclusion

Jigsaw, through its puzzle-piece-based visual design tool, lowers the technical barrier for using AI foundation models and expands designers' creative possibilities. The study validates the tool's effectiveness for non-technical designers and identifies potential areas for future improvement. Jigsaw's success demonstrates the significant potential of visual programming concepts in the integration of design and AI, opening new avenues for related research and applications.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147352/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3641920
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
2 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Human-LLM Collaboration, Prototyping & User Testing
work
Professions
UI/UX Designers, Product Designers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers