Detecting Input Recognition Errors and User Errors Using Gaze Dynamics in Virtual Reality

Eye Tracking & Gaze InteractionHuman Pose & Activity RecognitionImmersion & Presence ResearchUI/UX DesignersHCI Researchers

Document Title

Detecting Input Recognition Errors and User Errors using Gaze Dynamics in Virtual Reality

Document Information

  • Subject Area: Human-Computer Interaction (HCI), Virtual Reality (VR), Eye Tracking
  • Keywords: Input recognition errors, user errors, eye movement behavior, gaze dynamics, eye tracking, adaptive user interfaces

Research Background and Problem

  • Identified Problems/Challenges:

    • In virtual reality (VR) interactions, gesture recognition may lead to input recognition errors and user errors, significantly impacting user experience.
    • Current recognition systems struggle to completely eliminate these errors.
    • Error classification and detection have not been validated in multi-task scenarios, especially for naturally occurring errors.
  • Importance:

    • Timely identification and classification of errors can enhance system personalization and adaptability, reducing user interaction costs.
    • Accurate recognition of input recognition errors and user errors can improve interaction design and overall user experience.
  • Research Motivation and Related Work:

    • Previous studies (e.g., research by Peacock et al.) explored gaze-based binary classification models, but these focused on single tasks with artificially injected errors, lacking validation in complex tasks and natural errors.
    • This study aims to expand the application scope of prior models through cross-task validation and proposes a deep learning solution for three-class classification.

Solution

  • Proposed Method:

    • Develop a deep learning-based gaze tracking model to classify three types of input events based on users' natural eye movement behavior (gaze dynamics):
      • Intentional actions
      • Input recognition errors
      • User errors
    • Introduce cross-task consistency validation to ensure the model adapts to different task scenarios.
  • Innovations:

    • Proposed a Temporal Convolutional Network (TCN) for recognizing and classifying gaze features, applicable across multiple tasks.
    • Systematically compared the impact of injected and natural input errors on user gaze dynamics across different tasks and validated the model's generalization capability in cross-task and cross-user scenarios.
    • Leveraged gradient-driven eye movement features, avoiding the limitations of traditional task-specific feature selection, enabling data-driven dynamic pattern learning.
  • Implementation Steps:

    1. Collect gaze data from three tasks (Tile Search, Room Search, and Dice Game) and annotate input types.
    2. Preprocess raw data and extract features (including gaze velocity, discrete distribution points, etc.).
    3. Propose two classification hypotheses and validate them through quantitative experiments:
      • Hypothesis 1 (H1): Significant differences exist in time-series features for different types of input events.
      • Hypothesis 2 (H2): Features of the same type of events exhibit consistency across different tasks.
    4. Train and test classification models using TCN, evaluating performance across tasks and in cross-task scenarios.

Research Outcomes

  • Specific Results:

    1. Time-Series Feature Differences:
      • Eye movement features (e.g., gaze probability, saccade amplitude) of different input events (intentional actions, recognition errors, user errors) show significant differences in time-series data.
    2. Cross-Task Feature Consistency:
      • Despite variations in task control conditions (e.g., using handheld controllers vs gesture recognizers) and error types, gaze features remain consistent across all tasks.
      • Similar patterns can distinguish naturally occurring input recognition errors from artificially injected errors.
    3. Deep Learning Classification Performance:
      • In single-task classification experiments, the model achieved an AUC-ROC-OVR of 0.75 to 0.84.
      • In cross-task classification experiments, the combined model achieved an AUC-ROC-OVR of 0.78.
    4. Dimensionality Reduction Analysis:
      • Feature distributions in low-dimensional embedding spaces show clear clustering of the three input event types, supporting the algorithm's classification performance.
  • Experimental Evaluation Results:

    • Successfully achieved high-accuracy predictions of three event types using a unified model across the Tile Search, Room Search, and Dice Game tasks.
    • Experiments support the feasibility of using time-series features for classification and their independence from specific tasks.
  • Comparison with Existing Solutions:

    • Compared to previous binary classification models, this study supports three-class classification tasks and is applicable to more complex natural errors and multi-task scenarios.
    • The deep learning approach significantly improves robustness, avoiding the limitations of manual feature selection.
  • Limitations and Future Directions:

    • Limitations:
      • Although the classification model's accuracy (AUC-ROC-OVR=0.78) is significantly higher than random levels, higher performance is required for real-world system deployment.
      • The study is limited to point-and-click tasks; eye movement trends in other interaction modes (e.g., swiping, scrolling gestures) require further validation.
    • Future Directions:
      1. Explore the use of multimodal data (e.g., user hand movements) to optimize classification model performance.
      2. Develop real-time error detection systems and validate their impact on improving user interaction performance.
      3. Extend the study to more complex scenarios, such as multi-target selection or multi-gesture input environments.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/uist/85046/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3526113.3545628
At a Glance

Paper Snapshot

fact_check
dataset
Source
UIST
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Eye Tracking & Gaze Interaction, Human Pose & Activity Recognition, Immersion & Presence Research
work
Professions
UI/UX Designers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers