Gaze and Speech in Multimodal Human-Computer Interaction: A Scoping Review
Authors
Paper Title
Gaze and Speech in Multimodal Human-Computer Interaction: A Scoping Review
Publication Info
- Topic area: Multimodal human-computer interaction focusing on gaze and speech integration.
- Keywords: Multimodal interaction, gaze input, speech input, human-computer interaction, explicit interaction, implicit interaction, temporal alignment, fusion strategies, adaptive systems, privacy.
Background and Problem
- Problem / challenge: Despite decades of research, the integration of gaze and speech in human-computer interaction remains fragmented, with limited synthesis of recurring patterns, challenges, and opportunities.
- Significance: Gaze and speech offer complementary strengths—gaze provides spatial grounding, while speech conveys semantic richness—making their combination crucial for intuitive, adaptive, and hands-free interaction in emerging technologies like XR and AI glasses.
- Motivation and related work: Previous surveys have focused broadly on multimodal interaction or independently on gaze or speech, without analyzing their combined use. This review addresses the gap by systematically synthesizing 103 studies to reveal integration strategies, task domains, and challenges unique to gaze and speech pairing.
Solution
- Proposed approach: A scoping review categorizing gaze and speech interaction into explicit (intentional control signals) and implicit (behavioral cues) paradigms, analyzing their integration strategies, feature representations, and application domains.
- Novelty:
- First comprehensive synthesis of gaze and speech as a unified input channel.
- Detailed analysis of multimodal integration strategies and feature representations.
- Mapping of tasks and domains leveraging gaze and speech, highlighting design challenges and opportunities.
- Procedure and key techniques:
- Structured search across ACM-DL, IEEE Xplore, Scopus, and Web of Science using refined keyword queries.
- Screening and inclusion criteria focused on gaze and speech as input modalities in single-user HCI.
- Coding studies by interaction type, integration strategy, and task domains.
- Analysis of explicit and implicit interaction paradigms through functional roles, fusion strategies, and computational models.
Results
- Concrete findings:
- Explicit interaction (55 studies): Gaze provides spatial precision, while speech conveys commands or data. Common combinations include "Gaze to Point, Speech as Command" and "Gaze to Point + Confirm, Speech as Data."
- Implicit interaction (48 studies): Systems infer user states and intentions from natural gaze and speech behavior, using feature- or model-level fusion strategies.
- Temporal alignment models are critical; gaze often precedes speech by ~630 ms.
- Multimodal systems consistently outperform unimodal baselines in reducing ambiguity and enhancing expressiveness.
- Advantage over baselines: Multimodal integration improves robustness, expressiveness, and adaptability compared to gaze-only or speech-only systems.
- Experiments / evaluation:
- Studies evaluated across diverse tasks (e.g., selection, system control, reference resolution, affect recognition) using manual and automated annotation methods.
- Validation methods include cross-validation, K-fold splits, and limited cross-task or cross-user generalization.
- Limitations and future work:
- Temporal alignment challenges, inconsistent fusion strategies, and limited evaluation beyond controlled settings.
- Ethical concerns regarding privacy and transparency in gaze and speech monitoring.
- Need for broader evaluation across neurodiverse and socially constrained contexts.
Summary
This scoping review synthesizes 103 studies on gaze and speech integration in human-computer interaction, categorizing them into explicit and implicit paradigms. Explicit systems leverage gaze for spatial precision and speech for commands or data, while implicit systems infer user states and intentions from natural behavior. The review highlights recurring multimodal patterns, integration strategies, and challenges such as temporal alignment, fusion strategy selection, and privacy concerns. By consolidating fragmented research, it provides a foundation for designing robust, adaptive, and context-aware multimodal systems, particularly in emerging technologies like XR and AI glasses.
Research Questions / Practical Problems
Question signals indexed for this paper.
Based on Jaccard similarity of research subtopics & professions (≥60%)