Detecting Input Recognition Errors and User Errors Using Gaze Dynamics in Virtual Reality
Authors
Document Title
Detecting Input Recognition Errors and User Errors using Gaze Dynamics in Virtual Reality
Document Information
- Subject Area: Human-Computer Interaction (HCI), Virtual Reality (VR), Eye Tracking
- Keywords: Input recognition errors, user errors, eye movement behavior, gaze dynamics, eye tracking, adaptive user interfaces
Research Background and Problem
-
Identified Problems/Challenges:
- In virtual reality (VR) interactions, gesture recognition may lead to input recognition errors and user errors, significantly impacting user experience.
- Current recognition systems struggle to completely eliminate these errors.
- Error classification and detection have not been validated in multi-task scenarios, especially for naturally occurring errors.
-
Importance:
- Timely identification and classification of errors can enhance system personalization and adaptability, reducing user interaction costs.
- Accurate recognition of input recognition errors and user errors can improve interaction design and overall user experience.
-
Research Motivation and Related Work:
- Previous studies (e.g., research by Peacock et al.) explored gaze-based binary classification models, but these focused on single tasks with artificially injected errors, lacking validation in complex tasks and natural errors.
- This study aims to expand the application scope of prior models through cross-task validation and proposes a deep learning solution for three-class classification.
Solution
-
Proposed Method:
- Develop a deep learning-based gaze tracking model to classify three types of input events based on users' natural eye movement behavior (gaze dynamics):
- Intentional actions
- Input recognition errors
- User errors
- Introduce cross-task consistency validation to ensure the model adapts to different task scenarios.
- Develop a deep learning-based gaze tracking model to classify three types of input events based on users' natural eye movement behavior (gaze dynamics):
-
Innovations:
- Proposed a Temporal Convolutional Network (TCN) for recognizing and classifying gaze features, applicable across multiple tasks.
- Systematically compared the impact of injected and natural input errors on user gaze dynamics across different tasks and validated the model's generalization capability in cross-task and cross-user scenarios.
- Leveraged gradient-driven eye movement features, avoiding the limitations of traditional task-specific feature selection, enabling data-driven dynamic pattern learning.
-
Implementation Steps:
- Collect gaze data from three tasks (Tile Search, Room Search, and Dice Game) and annotate input types.
- Preprocess raw data and extract features (including gaze velocity, discrete distribution points, etc.).
- Propose two classification hypotheses and validate them through quantitative experiments:
- Hypothesis 1 (H1): Significant differences exist in time-series features for different types of input events.
- Hypothesis 2 (H2): Features of the same type of events exhibit consistency across different tasks.
- Train and test classification models using TCN, evaluating performance across tasks and in cross-task scenarios.
Research Outcomes
-
Specific Results:
- Time-Series Feature Differences:
- Eye movement features (e.g., gaze probability, saccade amplitude) of different input events (intentional actions, recognition errors, user errors) show significant differences in time-series data.
- Cross-Task Feature Consistency:
- Despite variations in task control conditions (e.g., using handheld controllers vs gesture recognizers) and error types, gaze features remain consistent across all tasks.
- Similar patterns can distinguish naturally occurring input recognition errors from artificially injected errors.
- Deep Learning Classification Performance:
- In single-task classification experiments, the model achieved an AUC-ROC-OVR of 0.75 to 0.84.
- In cross-task classification experiments, the combined model achieved an AUC-ROC-OVR of 0.78.
- Dimensionality Reduction Analysis:
- Feature distributions in low-dimensional embedding spaces show clear clustering of the three input event types, supporting the algorithm's classification performance.
- Time-Series Feature Differences:
-
Experimental Evaluation Results:
- Successfully achieved high-accuracy predictions of three event types using a unified model across the Tile Search, Room Search, and Dice Game tasks.
- Experiments support the feasibility of using time-series features for classification and their independence from specific tasks.
-
Comparison with Existing Solutions:
- Compared to previous binary classification models, this study supports three-class classification tasks and is applicable to more complex natural errors and multi-task scenarios.
- The deep learning approach significantly improves robustness, avoiding the limitations of manual feature selection.
-
Limitations and Future Directions:
- Limitations:
- Although the classification model's accuracy (AUC-ROC-OVR=0.78) is significantly higher than random levels, higher performance is required for real-world system deployment.
- The study is limited to point-and-click tasks; eye movement trends in other interaction modes (e.g., swiping, scrolling gestures) require further validation.
- Future Directions:
- Explore the use of multimodal data (e.g., user hand movements) to optimize classification model performance.
- Develop real-time error detection systems and validate their impact on improving user interaction performance.
- Extend the study to more complex scenarios, such as multi-target selection or multi-gesture input environments.
- Limitations:
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- In VR, how do users' natural gaze dynamics differ across input event types (intentional behavior, recognition errors, user errors)?Category: Time Series Semantic Retrieval and Trend AnalysisSimilar questionsarrow_forward
- Are temporal sequence features of gaze dynamics consistent across tasks and input events?Category: Time Series Semantic Retrieval and Trend AnalysisSimilar questionsarrow_forward
- Can deep learning models effectively classify input event types in cross-task scenarios?Category: Time Series Semantic Retrieval and Trend AnalysisSimilar questionsarrow_forward
Practical Problems
1- In VR, recognition errors and user errors degrade interaction experience, and existing systems struggle to quickly classify error types.Category: Time Series Semantic Retrieval and Trend AnalysisSimilar questionsarrow_forward
- 80%
Physical Keyboards in Virtual Reality: Analysis of Typing Performance and Effects of Avatar Hands
CHI '18· Eye Tracking & Gaze Interaction +1
- 80%
FocusFlow: 3D Gaze-Depth Interaction in Virtual Reality Leveraging Active Visual Depth Manipulation
CHI '24· Eye Tracking & Gaze Interaction +1
- 80%
DEEP: 3D Gaze Pointing in Virtual Reality Leveraging Eyelid Movement
UIST '22· Eye Tracking & Gaze Interaction +1
- 67%
Head-Coupled Kinematic Template Matching: A Prediction Model for Ray Pointing in VR
CHI '20· Eye Tracking & Gaze Interaction +2
- 67%
Searching Through Complex Worlds: Visual Search and Spatial Regularity Memory in Mixed Reality
CHI '26· Immersion & Presence Research +2
- 67%
The Eye–Head Mover Spectrum: Modelling Individual and Population Head Movement Tendencies in Virtual Reality
CHI '26· Immersion & Presence Research +2
- 67%
Less is More! Visual Suppression for Bottom-up and Top-down Attention in Dynamic Environments
CHI '26· Immersion & Presence Research +2
- 60%
An Explanation of Fitts' Law-like Performance in Gaze-Based Selection Tasks Using a Psychophysics Approach
CHI '19· Eye Tracking & Gaze Interaction
- 60%
Faces of Focus: A Study on the Facial Cues of Attentional States
CHI '20· Eye Tracking & Gaze Interaction +1
- 60%
RetroSketch: A Retrospective Method for Measuring Emotions and Presence in Virtual Reality
CHI '25· Immersion & Presence Research
Based on Jaccard similarity of research subtopics & professions (≥60%)