VoiceRay: 3D Object Selection Technique for Occluded and Dense Environments in Virtual Reality
Authors
Paper Title
VoiceRay: 3D Object Selection Technique for Occluded and Dense Environments in Virtual Reality
Publication Info
- Topic area: 3D object selection in dense and occluded virtual reality environments.
- Keywords: 3D object selection, virtual reality, raycasting, voice interaction, dense environments, occlusion, usability, cognitive load, presence, disambiguation techniques.
Background and Problem
- Problem / challenge: Existing 3D object selection techniques in VR, such as raycasting, struggle with disambiguation in dense and occluded environments. These methods often require multi-step interactions, scene modifications, or additional inputs, which can increase cognitive load, disrupt presence, and reduce usability.
- Significance: Effective and intuitive object selection is critical for VR applications, including training, medical simulations, and data visualization, where dense and occluded environments are common.
- Motivation and related work: Prior techniques include manual (e.g., AlphaCursor, LassoGrid), heuristic (e.g., BubbleRay), and behavioral approaches, each with trade-offs in accuracy, usability, and presence. Voice-based methods have shown promise but face challenges such as speech processing delays, memorization demands, and limited feedback. This paper addresses these gaps by integrating voice input with raycasting for efficient disambiguation.
Solution
- Proposed approach: VoiceRay, a voice-based raycasting technique that allows users to specify the ordinal position of a target along a ray (e.g., "second object") for disambiguation without altering the scene or requiring additional inputs.
- Novelty:
- Combines voice input with raycasting for 3D selection in dense and occluded VR environments.
- Avoids scene modifications, preserving user presence and immersion.
- Provides immediate visual feedback and low-latency speech recognition for efficient interaction.
- Systematic comparison with five existing techniques, analyzing performance, usability, cognitive load, and presence.
- Procedure and key techniques:
- Users point a ray at the target cluster using a VR controller.
- All intersected objects are outlined, and users speak the ordinal position of the intended target.
- The system transcribes the spoken command and highlights the selected object for confirmation.
- Speech recognition is implemented using Meta Voice SDK with a recognition delay of 120–150 ms.
Results
- Concrete findings:
- VoiceRay achieved the fastest selection time (significantly faster than BubbleRay and Raycasting) and one of the lowest error rates.
- SUS score for VoiceRay: 84.27 (Grade A, Excellent).
- Participants rated VoiceRay as less mentally demanding and less frustrating compared to other techniques.
- Advantage over baselines:
- Faster selection and lower error rates compared to Raycasting and BubbleRay.
- Higher usability and preference ratings than AlphaCursor, LassoGrid, and RayCursor.
- Preserved presence and realism compared to scene-altering techniques like AlphaCursor and LassoGrid.
- Experiments / evaluation:
- 24 participants tested six techniques (VoiceRay, AlphaCursor, RayCursor, LassoGrid, BubbleRay, Raycasting) in a dense VR environment with 220 spheres.
- Metrics: selection time, error rate, usability (SUS), cognitive load (NASA-TLX), and presence (IPQ).
- Follow-up pilot study explored alternative workflows (e.g., ButtonRay, VoiceLockRay) with no significant performance improvements.
- Limitations and future work:
- Limited to English-speaking participants; cross-linguistic robustness needs exploration.
- Tested with up to four intersected objects along a ray; future studies should investigate higher densities.
- Speech recognition may face challenges in noisy or multi-user environments, requiring further refinement.
Summary
VoiceRay is a voice-based raycasting technique designed for 3D object selection in dense and occluded VR environments. It enables faster and more accurate disambiguation compared to existing methods while preserving user presence and minimizing cognitive load. The study demonstrates its effectiveness through quantitative and qualitative evaluations, highlighting its potential for applications in VR training, medical simulations, and data visualization. Future work could extend VoiceRay to support more complex scenarios, larger object densities, and diverse user populations.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)