Jigsaw: Supporting Designers to Prototype Multimodal Applications by Chaining AI Foundation Models
Document Title
Jigsaw: Supporting Designers to Prototype Multimodal Applications by Assembling AI Foundation Models
Document Information
- Domain: Human-Computer Interaction, Design and AI Integration, Visual Programming
- Keywords: Prototyping, Machine Learning, Foundation Models, Multimodal, Visual Programming Interface
Research Background and Problem
-
Challenges:
- Designers have limited understanding of the full capabilities of AI foundation models, making it difficult to fully leverage their potential.
- Writing "AI-friendly" prompts and adjusting parameters is cumbersome.
- Integrating cross-platform, multimodal models is highly challenging.
- The prototyping process is slow, hindering effective rapid iteration and experimentation.
-
Significance: With the rapid development of AI foundation models, designers are attempting to integrate these models into creative workflows but face numerous technical and workflow barriers. Addressing these issues will enhance designers' efficiency and promote broader adoption of AI technologies.
-
Motivation: The authors aim to design a tool (Jigsaw) that allows designers to easily explore, combine, and utilize AI foundation models, thereby overcoming the aforementioned challenges.
Solution
-
Proposed Approach:
- Jigsaw uses puzzle pieces as a metaphor for foundation models, with each piece representing a function or model.
- Includes a "Catalog Panel" to display all available models, providing model descriptions, usage examples, and search functionality.
- Employs color-coded blocks and a snapping mechanism to clarify model compatibility, simplifying the combination of multimodal models.
- Integrates an "Assembly Assistant" feature to recommend model combinations based on natural language descriptions.
- Offers real-time feedback on intermediate results, enabling users to quickly observe and adjust during construction and debugging.
-
Innovations:
- An intuitive, puzzle-piece-based visual interface significantly lowers the technical barrier for designers using AI.
- Integrated prompt generation and parameter explanation features help users easily create effective AI model inputs.
- Seamless integration of multimodal models.
- A novel block-based visual programming interface tailored for non-technical designers.
-
Implementation Steps:
- Created a model library supporting text, images, video, 3D, audio, and sketches, comprising 39 models in total.
- Designed specific input, output types, and default parameters for each model.
- Implemented LLM-based "glue pieces" to support multi-model chaining tasks.
- Integrated real-time input and output panels, supporting multimodal input and output formats including text, images, and audio.
Research Outcomes
-
Key Results:
- Developed the Jigsaw system, simplifying the process for designers to use foundation models.
- Enabled users to construct complex AI-based design workflows through a puzzle-piece approach.
- Facilitated model exploration, rapid prototyping, and visualization of design records.
-
Advantages Over Existing Solutions:
- A block-based interface designed for ease of use by non-technical designers.
- Effectively guides designers in selecting suitable models and achieving model combinations.
- Supports cross-platform collaboration and rapid prototyping of multimodal models.
-
Experimental Results:
- Conducted user studies with ten designers, revealing that Jigsaw significantly helps designers:
- Explore new AI capabilities and gain deeper understanding of models through semantic search and puzzle-piece prompts.
- Rapidly construct and iterate AI creative workflows.
- Intuitively combine models across multiple tasks and modalities.
- Conducted user studies with ten designers, revealing that Jigsaw significantly helps designers:
-
Limitations and Future Directions:
- Currently, only one AI model is supported per task; future plans include expanding to multiple model options.
- Enhancing the Assembly Assistant's ability to handle complex tasks and support interactive iterations with designers.
- Expanding support for real-time video and audio stream input/output.
- Exploring the integration of multimodal large language model (MLLM) technology into Jigsaw.
- Creating a library of design templates to facilitate experience and workflow sharing among designers.
Conclusion
Jigsaw, through its puzzle-piece-based visual design tool, lowers the technical barrier for using AI foundation models and expands designers' creative possibilities. The study validates the tool's effectiveness for non-technical designers and identifies potential areas for future improvement. Jigsaw's success demonstrates the significant potential of visual programming concepts in the integration of design and AI, opening new avenues for related research and applications.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can designers efficiently explore, combine, and utilize AI foundation models through visualization tools?Category: Creative Workflows and Multi-Stage PipelinesSimilar questionsarrow_forward
- Can puzzle-style interfaces lower the technical barrier for non-technical designers using multimodal AI models?Category: Creative Workflows and Multi-Stage PipelinesSimilar questionsarrow_forward
- How can designers rapidly build and iterate AI creative workflows?Category: Creative Workflows and Multi-Stage PipelinesSimilar questionsarrow_forward
Practical Problems
1- Designers struggle to quickly build and iterate multimodal AI-based design workflows.Category: Creative Workflows and Multi-Stage PipelinesSimilar questionsarrow_forward
- 100%
GenTune: Toward Traceable Prompts to Improve Controllability of Image Refinement in Environment Design
UIST '25· Generative AI (Text, Image, Music, Video) +2
- 80%
Exploring Challenges and Opportunities to Support Designers in Learning to Co-create with AI-based Manufacturing Design Tools
CHI '23· Generative AI (Text, Image, Music, Video) +1
- 80%
AI-Assisted Causal Pathway Diagram for Human-Centered Design
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 80%
I-Card: A Generative AI-Supported Intelligent Design Method Card Deck
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 80%
Design Ideation with AI - Sketching, Thinking and Talking with Generative Machine Learning Models
DIS '23· Generative AI (Text, Image, Music, Video) +1
- 80%
PromptInfuser: How Tightly Coupling AI and UI Design Impacts Designers’ Workflows
DIS '24· Human-LLM Collaboration +1
- 67%
May AI? Design Ideation with Cooperative Contextual Bandits
CHI '19· Generative AI (Text, Image, Music, Video) +2
- 67%
ICONATE: Automatic Compound Icon Generation and Ideation
CHI '20· Generative AI (Text, Image, Music, Video) +2
- 67%
PlantoGraphy: Incorporating Iterative Design Process into Generative Artificial Intelligence for Landscape Rendering
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 67%
IEDS: Exploring an Intelli-Embodied Design Space Combining Designer, AR, and GAI to Support Industrial Conceptual Design
CHI '25· AR Navigation & Context Awareness +2
Based on Jaccard similarity of research subtopics & professions (≥60%)