VoiceRay: 3D Object Selection Technique for Occluded and Dense Environments in Virtual Reality

Social & Collaborative VRVoice User Interface (VUI) DesignEye Tracking & Gaze InteractionGame Developers & DesignersHCI Researchers

Paper Title

VoiceRay: 3D Object Selection Technique for Occluded and Dense Environments in Virtual Reality

Publication Info

  • Topic area: 3D object selection in dense and occluded virtual reality environments.
  • Keywords: 3D object selection, virtual reality, raycasting, voice interaction, dense environments, occlusion, usability, cognitive load, presence, disambiguation techniques.

Background and Problem

  • Problem / challenge: Existing 3D object selection techniques in VR, such as raycasting, struggle with disambiguation in dense and occluded environments. These methods often require multi-step interactions, scene modifications, or additional inputs, which can increase cognitive load, disrupt presence, and reduce usability.
  • Significance: Effective and intuitive object selection is critical for VR applications, including training, medical simulations, and data visualization, where dense and occluded environments are common.
  • Motivation and related work: Prior techniques include manual (e.g., AlphaCursor, LassoGrid), heuristic (e.g., BubbleRay), and behavioral approaches, each with trade-offs in accuracy, usability, and presence. Voice-based methods have shown promise but face challenges such as speech processing delays, memorization demands, and limited feedback. This paper addresses these gaps by integrating voice input with raycasting for efficient disambiguation.

Solution

  • Proposed approach: VoiceRay, a voice-based raycasting technique that allows users to specify the ordinal position of a target along a ray (e.g., "second object") for disambiguation without altering the scene or requiring additional inputs.
  • Novelty:
    1. Combines voice input with raycasting for 3D selection in dense and occluded VR environments.
    2. Avoids scene modifications, preserving user presence and immersion.
    3. Provides immediate visual feedback and low-latency speech recognition for efficient interaction.
    4. Systematic comparison with five existing techniques, analyzing performance, usability, cognitive load, and presence.
  • Procedure and key techniques:
    1. Users point a ray at the target cluster using a VR controller.
    2. All intersected objects are outlined, and users speak the ordinal position of the intended target.
    3. The system transcribes the spoken command and highlights the selected object for confirmation.
    4. Speech recognition is implemented using Meta Voice SDK with a recognition delay of 120–150 ms.

Results

  • Concrete findings:
    • VoiceRay achieved the fastest selection time (significantly faster than BubbleRay and Raycasting) and one of the lowest error rates.
    • SUS score for VoiceRay: 84.27 (Grade A, Excellent).
    • Participants rated VoiceRay as less mentally demanding and less frustrating compared to other techniques.
  • Advantage over baselines:
    • Faster selection and lower error rates compared to Raycasting and BubbleRay.
    • Higher usability and preference ratings than AlphaCursor, LassoGrid, and RayCursor.
    • Preserved presence and realism compared to scene-altering techniques like AlphaCursor and LassoGrid.
  • Experiments / evaluation:
    • 24 participants tested six techniques (VoiceRay, AlphaCursor, RayCursor, LassoGrid, BubbleRay, Raycasting) in a dense VR environment with 220 spheres.
    • Metrics: selection time, error rate, usability (SUS), cognitive load (NASA-TLX), and presence (IPQ).
    • Follow-up pilot study explored alternative workflows (e.g., ButtonRay, VoiceLockRay) with no significant performance improvements.
  • Limitations and future work:
    • Limited to English-speaking participants; cross-linguistic robustness needs exploration.
    • Tested with up to four intersected objects along a ray; future studies should investigate higher densities.
    • Speech recognition may face challenges in noisy or multi-user environments, requiring further refinement.

Summary

VoiceRay is a voice-based raycasting technique designed for 3D object selection in dense and occluded VR environments. It enables faster and more accurate disambiguation compared to existing methods while preserving user presence and minimizing cognitive load. The study demonstrates its effectiveness through quantitative and qualitative evaluations, highlighting its potential for applications in VR training, medical simulations, and data visualization. Future work could extend VoiceRay to support more complex scenarios, larger object densities, and diverse user populations.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222188/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790399
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Social & Collaborative VR, Voice User Interface (VUI) Design, Eye Tracking & Gaze Interaction
work
Professions
Game Developers & Designers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers