Training Person-Specific Gaze Estimators from User Interactions with Multiple Devices
Authors
Learning-based gaze estimation has significant potential to enable attentive user interfaces and gaze-based interaction on the billions of camera-equipped handheld devices and ambient displays. While training accurate person- and device-independent gaze estimators remains challenging, person-specific training is feasible but requires tedious data collection for each target device. To address these limitations, we present the first method to train person-specific gaze estimators across multiple devices. At the core of our method is a single convolutional neural network with shared feature extraction layers and device-specific branches that we train from face images and corresponding on-screen gaze locations. Detailed evaluations on a new dataset of interactions with five common devices (mobile phone, tablet, laptop, desktop computer, smart TV) and three common applications (mobile game, text editing, media center) demonstrate the significant potential of cross-device training. We further explore training with gaze locations derived from natural interactions, such as mouse or touch input.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 80%
HiFiGaze: Improving Eye Tracking Accuracy Using Screen Content Knowledge
CHI '26· Eye Tracking & Gaze Interaction +1
- 60%
Improving Discoverability and Expert Performance in Force-Sensitive Text Selection for Touch Devices with Mode Gauges
CHI '18· Force Feedback & Pseudo-Haptic Weight +1
- 60%
Pinpointing: Precise Head- and Eye-Based Target Selection for Augmented Reality
CHI '18· Eye Tracking & Gaze Interaction +1
- 60%
Doppio: Tracking UI Flows and Code Changes for App Development
CHI '18· Knowledge Worker Tools & Workflows +1
- 60%
Computational Support for Functionality Selection in Interaction Design
CHI '18· User Research Methods (Interviews, Surveys, Observation) +1
- 60%
Crowdsourcing Interface Feature Design with Bayesian Optimization
CHI '19· Crowdsourcing Task Design & Quality Control +1
- 60%
HandSee: Enabling Full Hand Interaction on Smartphone with Front Camera-based Stereo Vision
CHI '19· Hand Gesture Recognition +1
- 60%
GazeConduits: Calibration-Free Cross-Device Collaboration through Gaze and Touch
CHI '20· Eye Tracking & Gaze Interaction +1
- 60%
TapNet: The Design, Training, Implementation, and Applications of a Multi-Task Learning CNN for Off-Screen Mobile Input
CHI '21· Foot & Wrist Interaction +1
- 60%
Varv: Reprogrammable Interactive Software as a Declarative Data Structure
CHI '22· Prototyping & User Testing +1
Based on Jaccard similarity of research subtopics & professions (≥60%)