JustShape: Exploring Co-Speech Gestures for Multimodal LLM-Powered 3D Parametric Modeling
Honorable MentionAuthors
Paper Title
JustShape: Exploring Co-Speech Gestures for Multimodal LLM-Powered 3D Parametric Modeling
Publication Info
- Topic area: Multimodal interaction paradigms for 3D parametric modeling using large language models.
- Keywords: Co-speech gestures, multimodal interaction, parametric modeling, augmented reality, large language models, generative design, CAD, gesture recognition, user studies, design intent.
Background and Problem
- Problem / challenge: Traditional parametric modeling methods require significant technical expertise and create a "gulf of execution" between user intent and system operations. Text-only approaches lack the precision needed for parametric modeling, while gesture-based systems often focus on command substitution rather than intent-driven design.
- Significance: Lowering barriers to parametric modeling can democratize access to powerful design tools, enabling novices and embodied creators to engage in 3D modeling tasks without extensive training.
- Motivation and related work: Prior research explored GUI, programming-based, and sketch-based modeling, but these methods demand structured operations and technical proficiency. Multimodal LLMs have shown promise in bridging the "gulf of execution" in other domains but face challenges in conveying precise geometric and spatial attributes required for parametric modeling.
Solution
- Proposed approach: JustShape—a multimodal LLM-powered system that integrates co-speech gestures and speech inputs to generate parametric 3D models in augmented reality.
- Novelty:
- Elicitation study identifying the synergy between speech and co-speech gestures for expressing design intent.
- Multimodal fusion pipeline that parametrizes gestures and synthesizes them with speech for precise modeling.
- AR-based prototype enabling intuitive interaction and iterative refinement of parametric models.
- Empirical evaluation showing significant improvements in geometric accuracy and user experience compared to baseline methods.
- Procedure and key techniques:
- Conducted elicitation study to categorize co-speech gestures and their synergy with speech.
- Developed a multimodal fusion pipeline using gesture parametrization functions and a tool-augmented LLM.
- Implemented iterative design workflows in an AR interface, supporting real-time visualization, gesture tracking, and parametric editing.
- Evaluated system performance through user studies comparing co-speech gestures against speech-only and sketch+speech inputs.
Results
- Concrete findings:
- Co-speech gestures improved geometric accuracy by 40.1% compared to speech-only inputs and reduced interaction time by 47.3% compared to sketch+speech inputs.
- Gesture classification accuracy ranged from 82.2% to 91.8% across six shape attributes, with trajectory-based gestures achieving the highest parameter accuracy (81.4%).
- Multimodal fusion pipeline significantly outperformed baseline systems in prompt quality (53.73 vs. 32.37) and geometric fidelity (Chamfer Distance: 35.21 vs. 70.64).
- Advantage over baselines:
- Co-speech gestures demonstrated superior usability, enjoyment, and intuitiveness compared to sketch+speech and speech-only methods.
- Reduced cognitive load and mental demand, enabling novice designers to express design intent more effectively.
- Experiments / evaluation:
- User Study 1: Compared co-speech gestures, speech-only, and sketch+speech inputs across five fundamental CAD operations.
- User Study 2: Evaluated system usability in closed-ended and open-ended modeling tasks, with SUS score of 80.16 indicating "good" to "excellent" usability.
- Limitations and future work:
- Current system supports intermediate complexity models but lacks capabilities for industry-level designs.
- Gesture input and AR systems may cause fatigue or discomfort, highlighting the need for hardware improvements.
- Future work includes cross-platform deployment, integration with expert workflows, and scaling multimodal inputs to broader 3D modeling tasks.
Summary
JustShape introduces a multimodal interaction paradigm for parametric modeling, leveraging co-speech gestures and speech inputs to reduce barriers for novice users. The system integrates gesture parametrization and a multimodal fusion pipeline within an AR interface, enabling intuitive design workflows. User studies demonstrate significant improvements in geometric accuracy, usability, and user experience compared to traditional methods. While challenges remain in latency, robustness, and scalability, JustShape represents a promising step toward democratizing access to 3D parametric modeling and bridging embodied design practices with computational tools.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)