HiFiGaze: Improving Eye Tracking Accuracy Using Screen Content Knowledge
Honorable MentionAuthors
Paper Title
HiFiGaze: Improving Eye Tracking Accuracy Using Screen Content Knowledge
Publication Info
- Topic area: Gaze estimation on consumer computing devices using screen content knowledge.
- Keywords: Eye tracking, gaze estimation, screen reflections, consumer devices, high-resolution cameras, screen content knowledge, calibration-free, mobile devices, corneal reflection, user-facing cameras.
Background and Problem
- Problem / challenge: Conventional gaze tracking methods on commodity devices rely on appearance-based models, which provide limited accuracy (~2 cm mean error) and struggle with diverse screen content and lighting conditions.
- Significance: Accurate gaze tracking on everyday devices could enable widespread applications in accessibility, attention sensing, and gaze-aware interfaces without requiring specialized hardware.
- Motivation and related work: Prior methods, such as ScreenGlint, used screen reflections but required impractical screen conditions (e.g., full white fields). Advances in user-facing camera resolution (e.g., 4K) now make it feasible to capture detailed reflections of screen content in the eye, but leveraging this information robustly remains a challenge.
Solution
- Proposed approach: HiFiGaze, a gaze estimation method that combines high-resolution eye images with screen content knowledge to segment screen reflections and predict gaze points.
- Novelty:
- Utilizes screen content knowledge to segment and analyze screen reflections in the eye.
- Introduces multiple encoding strategies, including reflection vectors and heatmaps, to improve gaze prediction accuracy.
- Demonstrates calibration-free gaze tracking with improved accuracy on commodity devices.
- Explores the impact of camera placement (top vs. bottom) on gaze estimation performance.
- Procedure and key techniques:
- Preprocess eye images using MediaPipe landmarks and custom iris center correction (GrabCut, RANSAC, circle fitting).
- Encode screen content as thumbnails and correlate with segmented eye reflections to extract reflection heatmaps and vectors.
- Train a deep learning model with concatenated features (eye crops, screen thumbnails, reflection heatmaps/vectors) to predict gaze points.
- Evaluate performance across diverse screen content and lighting conditions, and assess the impact of camera placement.
Results
- Concrete findings:
- HiFiGaze reduces mean gaze error by ~18% (2.00 to 1.64 cm) compared to a baseline appearance-based model.
- Models incorporating screen content knowledge outperform baselines even under challenging conditions (e.g., dark screens).
- Bottom-camera placement improves accuracy by up to 18% for screen-reflection-based models.
- Advantage over baselines:
- Significant accuracy improvements over conventional appearance-based methods, particularly in challenging screen regions and lighting conditions.
- Reflection heatmap data alone performs comparably to baseline methods, offering potential privacy benefits.
- Experiments / evaluation:
- User study with 22 participants (~177K gaze instances) under natural conditions (sitting, standing, varying screen brightness).
- Supplemental study with 10 participants comparing top vs. bottom camera placement.
- Evaluated models using leave-one-participant-out cross-validation.
- Limitations and future work:
- Limited diversity in participant demographics (e.g., iris color, glasses wearers).
- Challenges with dark screen content and occlusions from eyelids/eyelashes.
- Need for broader data collection across dynamic conditions (e.g., walking, varied lighting) and device types (e.g., laptops, TVs).
- Potential improvements in iris center detection and adaptive modeling for physiological variability.
Summary
HiFiGaze introduces a novel gaze estimation method leveraging screen content knowledge and high-resolution camera imagery to segment screen reflections in the eye. This approach improves accuracy by ~18% over traditional appearance-based methods, achieving calibration-free gaze tracking on commodity devices. Experiments demonstrate robustness across diverse screen content and lighting conditions, with further accuracy gains from bottom-camera placement. Future work will address broader ecological validity, physiological variability, and dynamic usage scenarios, paving the way for practical, widely deployable gaze-aware applications.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 80%
Training Person-Specific Gaze Estimators from User Interactions with Multiple Devices
CHI '18· Eye Tracking & Gaze Interaction +1
- 80%
Dynamics of eye-hand coordination are flexibly preserved in eye-cursor coordination during an online, digital, object interaction task
CHI '23· Eye Tracking & Gaze Interaction +1
- 67%
A Simulation Model of Intermittently Controlled Point-and-Click Behaviour
CHI '21· Eye Tracking & Gaze Interaction +2
- 67%
In-Depth Mouse: Integrating Desktop Mouse into Virtual Reality
CHI '22· Eye Tracking & Gaze Interaction +1
- 67%
Graph4GUI: Graph Neural Networks for Representing Graphical User Interfaces
CHI '24· 360° Video & Panoramic Content +1
- 60%
Computational Interaction: Theory and Practice
CHI '18· Computational Methods in HCI
- 60%
An Explanation of Fitts' Law-like Performance in Gaze-Based Selection Tasks Using a Psychophysics Approach
CHI '19· Eye Tracking & Gaze Interaction
- 60%
Relative Design Acquisition: A Computational Approach for Creating Visual Interfaces to Steer User Choices
CHI '23· Computational Methods in HCI
- 60%
Model-based Evaluation of Recall-based Interaction Techniques
CHI '24· Computational Methods in HCI
- 60%
DeepSI: Interactive Deep Learning for Semantic Interaction
IUI '21· Computational Methods in HCI
Based on Jaccard similarity of research subtopics & professions (≥60%)