Predicting Usability and UX based on Eye Movements: Identifying Cross-Stimuli Interaction Patterns with Machine Learning
Authors
Paper Title
Predicting Usability and UX based on Eye Movements: Identifying Cross-Stimuli Interaction Patterns with Machine Learning
Publication Info
- Topic area: Eye-tracking and machine learning for usability and user experience (UX) evaluation.
- Keywords: Eye-tracking, usability, user experience, machine learning, pragmatic quality, hedonic quality, cross-stimuli generalization, AttrakDiff, UEQ, saccade patterns.
Background and Problem
- Problem / challenge: Current usability and UX evaluations rely heavily on self-report questionnaires, which are subjective and require active user participation. The potential of eye-tracking data to predict usability and UX ratings remains underexplored, especially in terms of generalizability across different digital products.
- Significance: Automating UX and usability evaluation through eye-tracking could reduce reliance on subjective methods, enabling passive and scalable assessments of digital products.
- Motivation and related work: Prior studies have shown that eye-tracking data can predict specific user behaviors and mental states but often lack generalizability across stimuli or domains. Few studies have attempted to predict usability and UX ratings using validated questionnaires like AttrakDiff and UEQ, leaving a gap in understanding cross-stimuli interaction patterns.
Solution
- Proposed approach: Use machine learning models to predict pragmatic quality (PQ) and hedonic quality (HQ) ratings from eye-tracking data collected during interactions with six websites.
- Novelty:
- Integration of validated questionnaires (AttrakDiff and UEQ) as labels for machine learning models.
- Analysis of cross-stimuli generalizability of eye-tracking-based predictions.
- Identification of specific eye-movement patterns (e.g., saccades and transitions) linked to PQ and HQ.
- Evaluation of five machine learning models, including SVM and MLP, for predicting usability and UX.
- Procedure and key techniques:
- Data collection from 121 participants using Tobii Pro Fusion eye-trackers across six websites.
- Feature engineering to extract third- and fourth-order eye-movement metrics, including saccade transitions and AOI-based features.
- Training and evaluation of machine learning models (SVM, MLP, KNN, RF, GB) using stratified 10-fold cross-validation and holdout validation.
- Analysis of feature importance and the impact of class imbalance on prediction performance.
Results
- Concrete findings:
- Within-stimulus MCC scores for SVM and MLP reached 0.751 and 0.780, respectively, with small to medium negative effect sizes in holdout validation (-0.313 to -0.492 for PQ, -0.141 to -0.182 for HQ).
- Cross-stimuli prediction MCC scores dropped to 0.196 for PQ and 0.338 for HQ, indicating limited generalizability.
- Fourth-order metrics, particularly saccade transitions, were most predictive, with regression saccades linked to HQ and longer successive saccades linked to PQ.
- Advantage over baselines:
- SVM and MLP outperformed KNN, RF, and GB in both within-stimulus and cross-stimuli predictions, showing higher MCC scores and narrower confidence intervals.
- HQ predictions were more generalizable across stimuli than PQ predictions.
- Experiments / evaluation:
- Data collected from six websites across two domains (German drinking water providers and German Alpine Association).
- Labels derived from short versions of AttrakDiff and UEQ questionnaires.
- Evaluation metrics included MCC, Cohen’s d, and 90% confidence intervals.
- Limitations and future work:
- Limited dataset size (121 participants) and class imbalance affected model performance.
- PQ ratings were highly stimulus-specific, hindering cross-stimuli generalization.
- Future work should explore larger datasets, multi-class predictions, and alternative interaction data (e.g., mouse movements).
Summary
This study demonstrates that eye-tracking data can moderately predict usability (PQ) and UX (HQ) ratings using machine learning models, with SVM and MLP achieving the best performance. Fourth-order metrics, such as saccade transitions, were particularly predictive, with regression saccades linked to HQ and longer successive saccades linked to PQ. While within-stimulus predictions were reliable, cross-stimuli generalization was limited, especially for PQ. The findings highlight the potential of eye-tracking as a complementary tool for usability and UX evaluation but underscore the need for further research to enhance generalizability and scalability.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 63%
When Should Users Check? Modeling Confirmation Frequency in Multi-Step Agentic AI Tasks
CHI '26· AI-Assisted Decision-Making & Automation +2
- 63%
Trust Formation in AI Delegation: The Interplay of Explainability and Anthropomorphism
CHI '26· Explainable AI (XAI) +2
- 63%
Do You (Dis)agree With Me? Modelling Implicit User Disagreement in Human–AI Interaction Using Gaze Data
CHI '26· Eye Tracking & Gaze Interaction +2
- 63%
Natural Expression of a Machine Learning Model's Uncertainty Through Verbal and Non-Verbal Behavior of Intelligent Virtual Agents
UIST '24· Eye Tracking & Gaze Interaction +2
Based on Jaccard similarity of research subtopics & professions (≥60%)