Predicting Gaze-based Target Selection in Augmented Reality Headsets based on Eye and Head Endpoint Distributions
Authors
Eye Tracking & Gaze InteractionAR Navigation & Context Awareness
Title of the Paper
Predicting Gaze-based Target Selection in Augmented Reality Headsets based on Eye and Head Endpoint Distributions
Paper Information
- Domain: Modeling gaze-based target selection in augmented reality
- Keywords: augmented reality, target selection, selection modeling, eye input, error prediction, multimodal input, head-eye coordination
Research Background and Problem
- Problem or Challenge: Target selection in augmented reality (AR) systems is a fundamental interaction task. However, gaze-based selection often suffers from inaccuracies due to tracking device noise and the instability of user gaze points. Additionally, many existing methods compromise the visual consistency of the user interface or introduce complexity by adding extra steps.
- Significance: With the widespread adoption of AR devices (e.g., HoloLens), gaze-based input can serve as an alternative to gestures and controllers, especially when users’ hands are occupied. Improving the accuracy of gaze-based selection can enhance the naturalness of system interactions.
- Motivation and Related Work:
- Current research primarily focuses on selection modeling in 2D interfaces or endpoint distribution modeling for 1D/2D targets.
- In 3D environments, most existing work emphasizes hand and head-based selection, with limited focus on endpoint distribution modeling for gaze-based 3D target selection.
Solution
- Proposed Method or Solution:
- Two gaze-based target prediction models based on endpoint distributions are proposed:
- Unimodal Model: Based solely on the eye endpoint at the time of user selection.
- Multimodal Model: Combines eye and head endpoints.
- Data Source:
- User experiments were conducted to collect data, including selection time, trajectories, and head direction.
- Model Foundation:
- Bayesian theory is used to construct a probabilistic target prediction model.
- Eye and head endpoint distributions are assumed to follow bivariate Gaussian distributions.
- Two gaze-based target prediction models based on endpoint distributions are proposed:
- Innovations:
- The newly proposed multimodal model integrates gaze and head direction.
- The model employs simple linear regression to capture factors influencing gaze behavior, such as target width and distance.
- It improves target selection accuracy without altering the visual appearance of the interaction interface.
- Implementation Steps and Key Techniques:
- Data Collection Experiment: Core data on eye and head endpoints were collected through experiments with controlled target width, distance, and confirmation mechanisms (e.g., “blink” or “air tap”).
- Model Construction and Fitting: Statistical data on user selection behavior were analyzed using linear regression, and parameters for endpoint distributions were modeled.
- Evaluation Experiment: The prediction models were compared with traditional selection methods, such as Visual Boundary Criterion (VBC).
Research Outcomes
- Specific Results:
- The proposed prediction models improved target selection accuracy by approximately 61% for smaller targets (e.g., 0.5° width).
- Both prediction models (unimodal and multimodal) significantly outperformed traditional VBC methods, particularly in scenarios with small targets and long distances.
- While the unimodal and multimodal models performed similarly in most cases, the multimodal model showed slight advantages for more distant targets.
- Advantages Compared to Existing Methods:
- The prediction models demonstrate higher stability than traditional VBC methods, handling more complex scenarios (e.g., smaller targets, greater distances, or dispersed backgrounds).
- They do not alter the UI interface or require additional operational steps.
- Experimental and Validation Results:
- In comprehensive user experiments, the proposed models achieved an average prediction accuracy of 94%-99%.
- The multimodal model performed better in scenarios requiring larger head movements, offering support in complex interaction contexts.
- Limitations and Future Directions:
- Limitations in Target Types and Environments:
- Experiments were conducted with static circular targets in controlled environments. Future work should validate the models in complex scenarios (e.g., dynamic targets or complex backgrounds).
- Impact of Search Behavior:
- Current experiments used predictable target sequences. Future research should explore head and eye behavior in free-search scenarios.
- Device Dependency:
- The study relied on eye-tracking data from Microsoft HoloLens 2. Hardware limitations may affect result accuracy, and the models should be validated on devices with higher precision.
- Limitations in Target Types and Environments:
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How can gaze-based target selection in AR devices be predicted based on eye and head endpoint distributions?Category: XR Target Selection and Interface ControlSimilar questionsarrow_forward
- In 3D target selection, do multimodal models (combining gaze and head direction) outperform unimodal models (gaze endpoint only)?Category: XR Target Selection and Interface ControlSimilar questionsarrow_forward
- Which factors (e.g., target width and distance) affect prediction accuracy of gaze behavior?Category: XR Target Selection and Interface ControlSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Users have low accuracy in gaze-based target selection in AR, often failing due to device noise or unstable gaze points.Category: XR Target Selection and Interface ControlSimilar questionsarrow_forward
- 100%
Reading on Smart Glasses: The Effect of Text Position, Presentation Type and Walking
CHI '18· Eye Tracking & Gaze Interaction +1
- 67%
Radi-Eye: Hands-free Radial Interfaces for 3D Interaction using Gaze-activated Head-crossing
CHI '21· Eye Tracking & Gaze Interaction +2
- 67%
Gaze on the Go: Effect of Spatial Reference Frame on Visual Target Acquisition During Physical Locomotion in Extended Reality
CHI '24· Full-Body Interaction & Embodied Input +2
- 67%
A picture is worth a thousand words? Investigating the Impact of Image Aids in AR on Memory Recall for Everyday Tasks
IUI '25· Eye Tracking & Gaze Interaction +2
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581042
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Eye Tracking & Gaze Interaction, AR Navigation & Context Awareness
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
4 related papers