Eyes on Many: Evaluating Gaze, Hand, and Voice for Multi-Object Selection in Extended Reality

Social & Collaborative VRImmersion & Presence ResearchEye Tracking & Gaze InteractionHand Gesture RecognitionUI/UX DesignersHCI Researchers

Paper Title

Eyes on Many: Evaluating Gaze, Hand, and Voice for Multi-Object Selection in Extended Reality

Publication Info

  • Topic area: Multi-object selection techniques in extended reality (XR) using gaze, hand, and voice modalities.
  • Keywords: Multi-object selection, extended reality, gaze interaction, hand gestures, voice commands, mode-switching, subselection, usability, user study, XR interaction.

Background and Problem

  • Problem / challenge: Multi-object selection in XR is underexplored, particularly the impact of mode-switching and subselection techniques on user performance. Existing research has focused on small target sets or limited modalities, leaving gaps in understanding how these techniques scale for larger sets or interact with each other.
  • Significance: Multi-object selection is critical for efficient interaction in XR, enabling faster manipulation of multiple items. Understanding the trade-offs between techniques can inform better design for XR systems.
  • Motivation and related work: Prior studies have explored gaze, hand, and voice interactions individually or in limited combinations, often focusing on small target sets or specific modalities. Persistent and quasi-mode-switching approaches have been studied in 2D contexts but are less understood in XR. This paper builds on these studies to evaluate combinations of mode-switching and subselection techniques in XR for larger target sets.

Solution

  • Proposed approach: A systematic evaluation of four mode-switching techniques (SemiPinch, FullPinch, DoublePinch, Voice) and three subselection techniques (Gaze+Dwell, Gaze+Pinch, Gaze+Voice) for multi-object selection in XR.
  • Novelty:
    1. Comprehensive empirical analysis of mode-switching and subselection combinations in XR.
    2. Identification of DoublePinch as the most effective mode-switching technique, especially when paired with Gaze+Pinch.
    3. Demonstration that SemiPinch is not viable for larger target sets due to instability and fatigue.
  • Procedure and key techniques:
    • Mode-switching techniques:
      • SemiPinch: Quasi-mode requiring a specific finger distance (2–7 cm).
      • FullPinch: Quasi-mode requiring a full pinch (<2 cm).
      • DoublePinch: Persistent-mode triggered by two consecutive pinches within 350 ms.
      • Voice: Persistent-mode activated by predefined voice commands.
    • Subselection techniques:
      • Gaze+Dwell (gD): Selection via gaze fixation for 450 ms.
      • Gaze+Pinch (gP): Selection via gaze targeting and a full pinch.
      • Gaze+Voice (gV): Selection via gaze targeting and a voice command.
    • User study design:
      • 30 participants performed multi-selection tasks with 6, 8, or 10 targets in a controlled VR environment.
      • Metrics included task completion time (TCT), mode-switching time (MST), error rates, inverse efficiency, and subjective feedback.

Results

  • Concrete findings:
    • DoublePinch + Gaze+Pinch achieved the best performance with the lowest task completion time, error rate, and highest efficiency.
    • SemiPinch had the highest error rates, fatigue, and lowest efficiency, particularly as target counts increased.
    • Persistent modes (DoublePinch, Voice) outperformed quasi-modes (FullPinch, SemiPinch) in stability, usability, and user preference.
    • Gaze+Pinch was the fastest subselection technique, while Gaze+Dwell was slower but less error-prone.
    • Gaze+Voice was perceived as tedious due to repetitive vocal commands.
  • Advantage over baselines:
    • Persistent modes reduced mode-switching errors and fatigue compared to quasi-modes.
    • DoublePinch scaled better with larger target sets, maintaining low error rates and high efficiency.
  • Experiments / evaluation:
    • Within-subjects design with 108 trials per participant.
    • Metrics included TCT, MST, mode error, accidental subselection ratio, inverse efficiency, and subjective workload (NASA-TLX, SUS, Borg CR10).
    • Results analyzed using repeated measures ANOVA and post-hoc tests.
  • Limitations and future work:
    • Study focused on serial multi-selection in a controlled 2D grid layout, limiting generalizability to 3D or irregular layouts.
    • Deselection and parallel selection techniques were not investigated.
    • Future work should explore more naturalistic 3D environments, parallel selection, and deselection strategies.

Summary

This study systematically evaluated mode-switching and subselection techniques for multi-object selection in XR, revealing that Persistent modes (DoublePinch, Voice) outperform Quasi modes (FullPinch, SemiPinch) in stability, efficiency, and usability. DoublePinch paired with Gaze+Pinch delivered the best performance, while SemiPinch proved unsuitable for larger target sets due to instability and fatigue. The findings provide actionable design recommendations for XR systems, emphasizing the importance of mode-switching stability and the interplay between mode-switching and subselection techniques. Future work should extend these insights to more complex 3D environments and parallel selection scenarios.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223245/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790513
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Social & Collaborative VR, Immersion & Presence Research, Eye Tracking & Gaze Interaction, Hand Gesture Recognition
work
Professions
UI/UX Designers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers