Scene2Hap: Generating Scene-Wide Haptics for VR from Scene Context with Multimodal LLMs
Best PaperAuthors
Paper Title
Scene2Hap: Generating Scene-Wide Haptics for VR from Scene Context with Multimodal LLMs
Publication Info
- Topic area: Haptic feedback design for immersive virtual reality using multimodal large language models.
- Keywords: Haptics, VR, multimodal LLMs, vibrotactile feedback, physics-inspired rendering, semantic inference, vibration propagation, immersive experience, user studies, scene-wide context.
Background and Problem
- Problem / challenge: Designing haptic feedback for VR scenes is time-consuming and lacks scalability. Existing methods fail to leverage full semantic information of objects or consider physical context and relationships between objects.
- Significance: Haptic feedback enhances immersion and realism in VR, making it crucial for creating engaging virtual environments.
- Motivation and related work: Prior work on haptic design tools and machine learning-based haptic generation has focused on isolated objects or manual efforts, without addressing scene-wide haptic rendering or vibration propagation. Scene2Hap builds on these limitations by automating haptic design at scale using multimodal LLMs and physics-inspired modeling.
Solution
- Proposed approach: Scene2Hap, an LLM-centered system that automatically designs object-level vibrotactile feedback for entire VR scenes based on semantic attributes and physical contexts.
- Novelty:
- A system architecture combining semantic inference and physics-inspired modeling for scalable haptic design.
- LLM-based haptic inference to estimate semantic and material properties of objects from multimodal scene data.
- Physics-inspired haptic rendering to simulate vibration propagation and attenuation in real-time.
- Empirical validation through three user studies demonstrating improved haptic realism, spatial awareness, and user experience.
- Procedure and key techniques:
- LLM-Based Haptic Inference: Extracts multimodal data (images, object names, dimensions) and uses chained LLM components (Scene Analyzer, Object Analyzer, Material Property Estimator, Vibration Describer) to infer semantic and physical attributes.
- Audio Retrieval/Generation: Uses vibration descriptions to retrieve or generate audio signals, which are converted into vibrotactile signals.
- Physics-Inspired Haptic Rendering: Builds a contact graph in real-time to model vibration propagation and attenuation based on material properties and spatial relationships.
- Implementation: Developed in Unity3D with a client-server model for LLM processing (GPT-4o) and audio generation (AudioGen).
Results
- Concrete findings:
- LLM-based haptic inference achieved high accuracy in estimating object semantics, material properties, and vibration behavior (average ratings >4 on a 5-point scale for most objects).
- Physics-inspired haptic rendering improved usability (utility, causality, consistency, saliency), materiality, and spatial awareness, with attenuated propagation rated highest.
- End-to-end pipeline enhanced realism, immersion, presence, and satisfaction in full VR scenes.
- Advantage over baselines:
- Scene2Hap automated haptic design across entire scenes, outperforming manual and isolated-object approaches.
- Attenuated vibration propagation significantly improved spatial awareness and material perception compared to no or full propagation conditions.
- Experiments / evaluation:
- Study 1: Evaluated LLM inference accuracy using human ratings and comparison to literature data for material properties.
- Study 2: Assessed the impact of vibration propagation on usability, materiality, and spatial awareness across three simplified VR scenes.
- Study 3: Investigated user experience in a full VR scene with seven vibration sources and interactive elements.
- Limitations and future work:
- Limited object semantics (binary vibration behavior) and simplified physical models.
- Dependence on specific LLMs and external audio retrieval/generation methods.
- Future directions include richer object states, higher-fidelity physical modeling, and broader haptic experiences (e.g., texture, friction, symbolic feedback).
Summary
Scene2Hap introduces an LLM-centered system for scalable haptic design in VR, combining semantic inference and physics-inspired modeling to generate adaptive vibrotactile feedback for entire scenes. Three user studies validated its effectiveness in improving realism, spatial awareness, and user experience. By automating haptic rendering based on scene-wide context, Scene2Hap addresses key limitations of prior approaches and offers practical benefits for VR designers. Future work aims to expand its capabilities for richer and more diverse haptic experiences.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 63%
Roomify: Spatially-Grounded Style Transformation for Immersive Virtual Environments
CHI '26· Social & Collaborative VR +3
- 63%
RoboHaptics: Designing Haptic Interactions for Lower Body with Quadruped Robot Dogs
CHI '26· Mid-Air Haptics (Ultrasonic) +3
Based on Jaccard similarity of research subtopics & professions (≥60%)