GestuProp: 3D Virtual Reality Prop Generation with Co-Speech Gestures
Authors
Paper Title
GestuProp: 3D Virtual Reality Prop Generation with Co-Speech Gestures
Publication Info
- Topic area: Virtual Reality (VR) object creation using multimodal interaction.
- Keywords: Virtual Reality, 3D object generation, co-speech gestures, multimodal interaction, generative AI, user-generated content, gesture-based interaction, immersive environments, human-computer interaction, design communication.
Background and Problem
- Problem / challenge: Existing 3D object creation methods in VR, such as sculpting or modular composition, are complex and require expertise, limiting accessibility for ordinary users. Current generative AI approaches often rely on text or sketches, which are less natural in VR and fail to fully capture user intent.
- Significance: Simplifying VR object creation can democratize content generation, enhance user creativity, and improve engagement in applications like gaming, education, and design.
- Motivation and related work: Prior research has explored multimodal inputs like voice and gestures for VR tasks, but these are primarily command-based and not suited for open-ended object creation. Iconic gestures, which convey shape and spatial relationships, remain underexplored in this context. This paper addresses the gap by focusing on co-speech gestures for intuitive 3D object generation.
Solution
- Proposed approach: GestuProp, a VR system that enables users to generate 3D props through co-speech gestures, combining gestures for spatial constraints and speech for semantic details.
- Novelty:
- Introduction of a co-speech gesture–guided pipeline for 3D object creation in VR.
- Development of a proxy model system to balance real-time responsiveness with generation quality.
- Insights into gesture–speech coordination and its impact on user satisfaction and interaction patterns.
- Exploration of user-driven design considerations for multimodal VR systems.
- Procedure and key techniques:
- Conducted a formative study with 30 participants to analyze gesture–speech interactions and derive design requirements.
- Designed a system pipeline with four stages: input (gesture and speech), mapping (gesture–speech integration), proxy model generation, and final 3D model generation.
- Implemented proxy models for real-time feedback and layered generation using AI tools like Stable Diffusion and Turbo-v1.0.
- Evaluated the system through a user study with 14 participants, focusing on usability, gesture patterns, and satisfaction.
Results
- Concrete findings:
- GestuProp achieved a mean SUS usability score of 78.0, indicating good usability.
- Proxy model generation took 2–3 seconds, while final models required 7–15 seconds depending on complexity.
- Participants showed high satisfaction with object orientation for directional items (e.g., guns, swords) but lower satisfaction for ambiguous forms like toys.
- Advantage over baselines: GestuProp provides a natural and intuitive interaction method, reducing learning effort compared to traditional 3D modeling or text-based generation systems.
- Experiments / evaluation:
- Conducted a 2×7 within-subjects user study with 14 participants across two VR environments (neutral and tech-themed) and seven object categories (e.g., pens, swords, toys).
- Collected 2,352 satisfaction ratings and conducted thematic analysis of interviews.
- Found that gesture–speech coordination patterns (e.g., semantic division of labor, temporal sequencing) significantly influenced user experience.
- Limitations and future work:
- Limited gesture vocabulary and reliance on four basic gestures.
- AI generation models struggled with IP-specific or highly regularized objects.
- Environmental context had limited influence due to simplistic scene design.
- Small sample size (14 participants) limits generalizability; larger studies are needed.
Summary
GestuProp introduces a novel co-speech gesture–guided system for 3D object creation in VR, leveraging gestures for spatial constraints and speech for semantic input. The system demonstrated good usability and user satisfaction, with participants appreciating its natural interaction and creative potential. Key findings highlight the importance of gesture–speech coordination and user-driven design considerations. While limitations in gesture vocabulary and AI model capabilities remain, GestuProp lays the groundwork for democratizing VR content creation and supporting applications in gaming, design, and collaborative environments. Future work should expand gesture sets, refine AI capabilities, and explore richer environmental contexts.
Research Questions / Practical Problems
Question signals indexed for this paper.
Based on Jaccard similarity of research subtopics & professions (≥60%)