Scene2Hap: Generating Scene-Wide Haptics for VR from Scene Context with Multimodal LLMs

Best Paper
Mid-Air Haptics (Ultrasonic)Social & Collaborative VRImmersion & Presence ResearchGame Developers & DesignersUI/UX DesignersAI/ML Researchers & Engineers

Paper Title

Scene2Hap: Generating Scene-Wide Haptics for VR from Scene Context with Multimodal LLMs

Publication Info

  • Topic area: Haptic feedback design for immersive virtual reality using multimodal large language models.
  • Keywords: Haptics, VR, multimodal LLMs, vibrotactile feedback, physics-inspired rendering, semantic inference, vibration propagation, immersive experience, user studies, scene-wide context.

Background and Problem

  • Problem / challenge: Designing haptic feedback for VR scenes is time-consuming and lacks scalability. Existing methods fail to leverage full semantic information of objects or consider physical context and relationships between objects.
  • Significance: Haptic feedback enhances immersion and realism in VR, making it crucial for creating engaging virtual environments.
  • Motivation and related work: Prior work on haptic design tools and machine learning-based haptic generation has focused on isolated objects or manual efforts, without addressing scene-wide haptic rendering or vibration propagation. Scene2Hap builds on these limitations by automating haptic design at scale using multimodal LLMs and physics-inspired modeling.

Solution

  • Proposed approach: Scene2Hap, an LLM-centered system that automatically designs object-level vibrotactile feedback for entire VR scenes based on semantic attributes and physical contexts.
  • Novelty:
    1. A system architecture combining semantic inference and physics-inspired modeling for scalable haptic design.
    2. LLM-based haptic inference to estimate semantic and material properties of objects from multimodal scene data.
    3. Physics-inspired haptic rendering to simulate vibration propagation and attenuation in real-time.
    4. Empirical validation through three user studies demonstrating improved haptic realism, spatial awareness, and user experience.
  • Procedure and key techniques:
    • LLM-Based Haptic Inference: Extracts multimodal data (images, object names, dimensions) and uses chained LLM components (Scene Analyzer, Object Analyzer, Material Property Estimator, Vibration Describer) to infer semantic and physical attributes.
    • Audio Retrieval/Generation: Uses vibration descriptions to retrieve or generate audio signals, which are converted into vibrotactile signals.
    • Physics-Inspired Haptic Rendering: Builds a contact graph in real-time to model vibration propagation and attenuation based on material properties and spatial relationships.
    • Implementation: Developed in Unity3D with a client-server model for LLM processing (GPT-4o) and audio generation (AudioGen).

Results

  • Concrete findings:
    • LLM-based haptic inference achieved high accuracy in estimating object semantics, material properties, and vibration behavior (average ratings >4 on a 5-point scale for most objects).
    • Physics-inspired haptic rendering improved usability (utility, causality, consistency, saliency), materiality, and spatial awareness, with attenuated propagation rated highest.
    • End-to-end pipeline enhanced realism, immersion, presence, and satisfaction in full VR scenes.
  • Advantage over baselines:
    • Scene2Hap automated haptic design across entire scenes, outperforming manual and isolated-object approaches.
    • Attenuated vibration propagation significantly improved spatial awareness and material perception compared to no or full propagation conditions.
  • Experiments / evaluation:
    • Study 1: Evaluated LLM inference accuracy using human ratings and comparison to literature data for material properties.
    • Study 2: Assessed the impact of vibration propagation on usability, materiality, and spatial awareness across three simplified VR scenes.
    • Study 3: Investigated user experience in a full VR scene with seven vibration sources and interactive elements.
  • Limitations and future work:
    • Limited object semantics (binary vibration behavior) and simplified physical models.
    • Dependence on specific LLMs and external audio retrieval/generation methods.
    • Future directions include richer object states, higher-fidelity physical modeling, and broader haptic experiences (e.g., texture, friction, symbolic feedback).

Summary

Scene2Hap introduces an LLM-centered system for scalable haptic design in VR, combining semantic inference and physics-inspired modeling to generate adaptive vibrotactile feedback for entire scenes. Three user studies validated its effectiveness in improving realism, spatial awareness, and user experience. By automating haptic rendering based on scene-wide context, Scene2Hap addresses key limitations of prior approaches and offers practical benefits for VR designers. Future work aims to expand its capabilities for richer and more diverse haptic experiences.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222083/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791297
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
Best Paper
group
Authors
5 authors
sell
Subtopics
Mid-Air Haptics (Ultrasonic), Social & Collaborative VR, Immersion & Presence Research
work
Professions
Game Developers & Designers, UI/UX Designers, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
2 related papers