Title of the Paper

Predicting Gaze-based Target Selection in Augmented Reality Headsets based on Eye and Head Endpoint Distributions

Paper Information

  • Domain: Modeling gaze-based target selection in augmented reality
  • Keywords: augmented reality, target selection, selection modeling, eye input, error prediction, multimodal input, head-eye coordination

Research Background and Problem

  • Problem or Challenge: Target selection in augmented reality (AR) systems is a fundamental interaction task. However, gaze-based selection often suffers from inaccuracies due to tracking device noise and the instability of user gaze points. Additionally, many existing methods compromise the visual consistency of the user interface or introduce complexity by adding extra steps.
  • Significance: With the widespread adoption of AR devices (e.g., HoloLens), gaze-based input can serve as an alternative to gestures and controllers, especially when users’ hands are occupied. Improving the accuracy of gaze-based selection can enhance the naturalness of system interactions.
  • Motivation and Related Work:
    • Current research primarily focuses on selection modeling in 2D interfaces or endpoint distribution modeling for 1D/2D targets.
    • In 3D environments, most existing work emphasizes hand and head-based selection, with limited focus on endpoint distribution modeling for gaze-based 3D target selection.

Solution

  • Proposed Method or Solution:
    1. Two gaze-based target prediction models based on endpoint distributions are proposed:
      • Unimodal Model: Based solely on the eye endpoint at the time of user selection.
      • Multimodal Model: Combines eye and head endpoints.
    2. Data Source:
      • User experiments were conducted to collect data, including selection time, trajectories, and head direction.
    3. Model Foundation:
      • Bayesian theory is used to construct a probabilistic target prediction model.
      • Eye and head endpoint distributions are assumed to follow bivariate Gaussian distributions.
  • Innovations:
    • The newly proposed multimodal model integrates gaze and head direction.
    • The model employs simple linear regression to capture factors influencing gaze behavior, such as target width and distance.
    • It improves target selection accuracy without altering the visual appearance of the interaction interface.
  • Implementation Steps and Key Techniques:
    1. Data Collection Experiment: Core data on eye and head endpoints were collected through experiments with controlled target width, distance, and confirmation mechanisms (e.g., “blink” or “air tap”).
    2. Model Construction and Fitting: Statistical data on user selection behavior were analyzed using linear regression, and parameters for endpoint distributions were modeled.
    3. Evaluation Experiment: The prediction models were compared with traditional selection methods, such as Visual Boundary Criterion (VBC).

Research Outcomes

  • Specific Results:
    1. The proposed prediction models improved target selection accuracy by approximately 61% for smaller targets (e.g., 0.5° width).
    2. Both prediction models (unimodal and multimodal) significantly outperformed traditional VBC methods, particularly in scenarios with small targets and long distances.
    3. While the unimodal and multimodal models performed similarly in most cases, the multimodal model showed slight advantages for more distant targets.
  • Advantages Compared to Existing Methods:
    • The prediction models demonstrate higher stability than traditional VBC methods, handling more complex scenarios (e.g., smaller targets, greater distances, or dispersed backgrounds).
    • They do not alter the UI interface or require additional operational steps.
  • Experimental and Validation Results:
    1. In comprehensive user experiments, the proposed models achieved an average prediction accuracy of 94%-99%.
    2. The multimodal model performed better in scenarios requiring larger head movements, offering support in complex interaction contexts.
  • Limitations and Future Directions:
    1. Limitations in Target Types and Environments:
      • Experiments were conducted with static circular targets in controlled environments. Future work should validate the models in complex scenarios (e.g., dynamic targets or complex backgrounds).
    2. Impact of Search Behavior:
      • Current experiments used predictable target sequences. Future research should explore head and eye behavior in free-search scenarios.
    3. Device Dependency:
      • The study relied on eye-tracking data from Microsoft HoloLens 2. Hardware limitations may affect result accuracy, and the models should be validated on devices with higher precision.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/96575/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581042
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Eye Tracking & Gaze Interaction, AR Navigation & Context Awareness
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
4 related papers