DesignMinds: Enhancing Video-Based Design Ideation with a Vision-Language Model and a Context-Injected Large Language Model
Authors
Ideation is a critical component of video-based design (VBD), where videos serve as the primary medium for design exploration and inspiration. The emergence of generative AI offers considerable potential to enhance this process by streamlining video analysis and facilitating idea generation. In this paper, we present DesignMinds, a prototype that integrates a state-of-the-art Vision-Language Model (VLM) with a context-enhanced Large Language Model (LLM) to support ideation in VBD. To evaluate DesignMinds, we conducted a between-subject study with 35 design practitioners, comparing its performance to a baseline condition. Our results demonstrate that DesignMinds significantly enhances the flexibility and originality of ideation, while also increasing task engagement. Importantly, the introduction of this technology did not negatively impact user experience, technology acceptance, or usability.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
ICONATE: Automatic Compound Icon Generation and Ideation
CHI '20· Generative AI (Text, Image, Music, Video) +2
- 100%
”Clay to Play With”: Generative AI Tools in UX and Industrial Design Practice
DIS '24· Generative AI (Text, Image, Music, Video) +2
- 80%
Exploring Challenges and Opportunities to Support Designers in Learning to Co-create with AI-based Manufacturing Design Tools
CHI '23· Generative AI (Text, Image, Music, Video) +1
- 80%
Learning Personal Style from Few Examples
DIS '21· Generative AI (Text, Image, Music, Video) +1
- 80%
Design Ideation with AI - Sketching, Thinking and Talking with Generative Machine Learning Models
DIS '23· Generative AI (Text, Image, Music, Video) +1
- 67%
Vinci: An Intelligent Graphic Design System for Generating Advertising Posters
CHI '21· Generative AI (Text, Image, Music, Video) +1
- 67%
StyleMe: Towards Intelligent Fashion Generation with Designer Style
CHI '23· Generative AI (Text, Image, Music, Video) +1
- 67%
Jigsaw: Supporting Designers to Prototype Multimodal Applications by Chaining AI Foundation Models
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 67%
CreativeConnect: Supporting Reference Recombination for Graphic Design Ideation with Generative AI
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 67%
Dancing With Chains: Ideating Under Constraints With UIDEC in UI/UX Design
CHI '25· 360° Video & Panoramic Content +2
Based on Jaccard similarity of research subtopics & professions (≥60%)