Do You (Dis)agree With Me? Modelling Implicit User Disagreement in Human–AI Interaction Using Gaze Data
Authors
Paper Title
Do You (Dis)agree With Me? Modelling Implicit User Disagreement in Human–AI Interaction Using Gaze Data
Publication Info
- Topic area: Implicit disagreement detection in human–AI interaction using gaze and facial data.
- Keywords: Human–AI interaction, disagreement detection, gaze tracking, facial expressions, machine learning, personalised models, multimodal analysis, implicit feedback, cognitive-affective states, dataset release.
Background and Problem
- Problem / challenge: Current AI systems lack mechanisms to detect user disagreement implicitly, relying instead on explicit feedback, which can disrupt interaction and reduce user trust. Detecting disagreement from passive signals like gaze and facial expressions is challenging due to subtle and individualised responses.
- Significance: Implicit disagreement detection can enhance the responsiveness and trustworthiness of AI systems by enabling adaptive and non-intrusive interaction, reducing cognitive load on users.
- Motivation and related work: Prior studies have explored gaze and facial data for related tasks like emotion recognition and cognitive state monitoring but have not focused on implicit disagreement detection. This paper addresses this gap by investigating gaze and facial signals in a controlled human–AI interaction scenario.
Solution
- Proposed approach: A machine learning framework to detect implicit user disagreement using gaze and facial data collected during an image-captioning task.
- Novelty:
- Development of a publicly available dataset of gaze and facial data with binary agreement/disagreement annotations.
- Exploration of personalised vs generalised machine learning models for disagreement detection.
- Analysis of gaze-based feature selection and time-window selection for improving detection performance.
- Procedure and key techniques:
- Conducted a controlled user study with 30 participants evaluating image-caption pairs.
- Collected gaze data using a Tobii Pro Fusion eye tracker and facial data using a Luxonis OAK-D camera.
- Extracted gaze-based features (e.g., fixation count, saccade duration) and facial action units (AUs).
- Compared personalised and generalised models using classical machine learning algorithms.
- Investigated the effects of feature subsets and time-window selection on model performance.
Results
- Concrete findings:
- Personalised gaze-only models achieved a balanced accuracy of 68.40%, outperforming generalised models (57.00%).
- Multimodal models combining gaze and facial data did not improve performance, with balanced accuracy close to chance (≈51.75%).
- Gaze features like fixation counts, saccade dynamics, and pupil diameter variability were most predictive of disagreement.
- Time-window selection showed that disagreement-related signals often concentrated in the last few seconds before a decision.
- Advantage over baselines:
- Personalised models showed significant performance improvement (mean accuracy increase ≈7.18%) compared to generalised models.
- Gaze-only models outperformed multimodal and facial-only models, highlighting the utility of gaze data for disagreement detection.
- Experiments / evaluation:
- Conducted group-based and user-based 5-fold cross-validation.
- Evaluated the impact of feature subsets (FA, FB, FPool) and time windows (full recording, last 11 seconds, last 3 seconds).
- Used statistical tests (e.g., Wilcoxon signed-rank) to confirm the significance of personalised configurations.
- Limitations and future work:
- Dataset size (30 participants) limits generalisability; results are exploratory.
- Subtle stimuli (single-word caption errors) may not elicit strong disagreement signals.
- Facial data contributed little predictive value due to neutral expressions and coarse feature processing.
- Future work should explore richer stimuli, larger datasets, multimodal fusion techniques, and ethical considerations for real-world deployment.
Summary
This study investigates implicit disagreement detection in human–AI interaction using gaze and facial data. A dataset of 30 participants was collected during an image-captioning task, and machine learning models were trained to predict agreement/disagreement. Personalised gaze-only models achieved the highest performance (68.40% balanced accuracy), while multimodal and facial-only models underperformed. Feature selection and time-window analysis revealed that disagreement signals are highly individualised and often concentrated near decision points. The findings highlight the potential of gaze data for adaptive AI systems but underscore the need for richer stimuli, larger datasets, and ethical considerations for practical deployment. The dataset and insights are publicly available to support further research.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
Mind The Gap: Designers and Standards on Algorithmic System Transparency for Users
CHI '24· Explainable AI (XAI) +2
- 71%
Sensemaking in Multi-Agent LLM Interfaces: How Users Interpret Transparency and Trustworthiness Cues
CHI '26· Human-LLM Collaboration +2
- 71%
Trust Formation in AI Delegation: The Interplay of Explainability and Anthropomorphism
CHI '26· Explainable AI (XAI) +2
- 71%
How Much Trust is Enough? Towards Calibrating Trust in Technology
CHI '26· Explainable AI (XAI) +2
- 71%
When the Codec Hallucinates: User Perceptions of Miscompressed Images
CHI '26· Explainable AI (XAI) +2
- 71%
Sensemaking in User-Driven Algorithm Auditing: A Case Study on Gender Bias in an Image Captioning Model
CHI '26· Explainable AI (XAI) +2
- 71%
Good Accessibility, Handcuffed Creativity: AI-Generated UIs Between Accessibility Guidelines and Practitioners’ Expectations
DIS '25· Explainable AI (XAI) +2
- 71%
Natural Expression of a Machine Learning Model's Uncertainty Through Verbal and Non-Verbal Behavior of Intelligent Virtual Agents
UIST '24· Eye Tracking & Gaze Interaction +2
- 67%
Don’t Just Tell Me, Ask Me: AI Systems that Intelligently Frame Explanations as Questions Improve Human Logical Discernment Accuracy over Causal AI explanations
CHI '23· Explainable AI (XAI) +1
- 63%
Characterizing User-Reported Risks across LLM Chatbots
CHI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)