Integrating Gaze and Speech for Enabling Implicit Interactions
Authors
Document Title
Integrating Gaze and Speech for Enabling Implicit Interactions
Document Information
- Field: Human-Computer Interaction (HCI)
- Keywords: Implicit annotation, natural gaze, speech interface, natural language processing, semantic similarity
Research Background and Problem Statement
-
Challenges:
- Using a single input modality (e.g., speech or gaze) often suffers from limitations inherent to that modality, restricting the naturalness and convenience of human-computer interaction.
- Existing technologies primarily use speech for explicit input (e.g., clear voice commands), while the semantics and complexity of natural speech remain underutilized. Additionally, relying solely on gaze for annotation can be inefficient in certain scenarios, such as rapid text scanning.
- Current gaze-based implicit interaction techniques lack reliability when faced with more natural user behaviors, especially when there is no explicit mapping between gaze patterns and speech annotations.
-
Importance of the Research:
- Combining speech and gaze input modalities can better capture user intent and contextual information, enhancing the naturalness and accuracy of human-computer interaction.
- In digital reading environments or similar scenarios, precise anchoring of speech annotations to text is crucial for user experience and task efficiency.
-
Motivation and Related Work:
- Classic studies (e.g., Bolt's "Put-that-there") demonstrate that combining multiple input modalities can complement each other and overcome individual limitations.
- The integration of speech and gaze has been widely studied, but mostly in explicit input scenarios, with insufficient exploration of implicit interaction contexts.
- Current gaze-based techniques often perform poorly during rapid reading, while the rich semantic information in natural speech has not been integrated with gaze data for more complex interaction scenarios.
Solution
-
Proposed Method or Solution:
- A machine learning-based data processing pipeline is proposed, integrating users' natural gaze behavior with the semantic information of speech annotations (using BERT semantic embeddings) and the synchronization between speech and gaze, to accurately predict the text segment corresponding to speech annotations.
- The pipeline includes multi-stage feature extraction: combining gaze pattern features, speech semantic similarity features, and implicit referencing features.
-
Innovations:
- Tight integration of speech and gaze modalities to capture contextual information that cannot be obtained through a single modality.
- Novel feature extraction for implicit anchoring of speech annotations, including synchronization metrics between gaze and speech.
- Development of a multimodal classifier based on natural user behavior, enhancing robustness and adaptability to various complex forms of speech annotations.
-
Implementation Steps and Techniques:
- Data Preprocessing: Extract annotation segments from speech data and calculate relevant gaze patterns from gaze data, marking text segments as candidate regions.
- Feature Extraction:
- Use the Sentence-BERT model to capture semantic similarity between speech and text embeddings.
- Extract implicit referencing features, including gaze dwell time, count, minimum distance, etc.
- Construct predictive features based on users' global reading patterns.
- Model Training: Train a logistic regression model, integrating gaze and speech features, and perform cross-validation for both user-dependent and independent modes.
- Evaluation and Analysis: Test the model using F1-Score, Average Precision (AP), and analyze feature importance based on SHAP.
Research Outcomes
-
Specific Results:
- The proposed algorithm achieved an average F1-Score of 0.90 in the task of precise anchoring of speech annotations.
- By combining gaze and speech modalities, the classifier consistently improved performance across different types of annotations (content-related, reflective, implicit referencing).
- Demonstrated robustness in handling implicit speech annotations, such as during rapid scanning or implicit referencing scenarios, with a 12% improvement in average precision compared to existing gaze-only solutions.
-
Comparative Advantages:
- Compared to existing gaze-only implicit interaction solutions (e.g., GAVIN), the multimodal approach significantly improved classification accuracy, especially in cases of text proximity or complex natural speech expressions.
-
Experimental and Evaluation Results:
- Both user-dependent and independent models showed strong adaptability to various speech annotation scenarios.
- Analysis of speech transcription accuracy revealed that even with higher error rates in automatic speech recognition (ASR), the model outperformed gaze-only methods.
-
Limitations and Future Directions:
- Limitations:
- Recognition of implicit referencing speech relies on certain pronouns (e.g., "this"), which may fail in complex language patterns.
- Validation was limited to text anchoring, lacking exploration of other digital content (e.g., videos, images).
- Anchoring was not refined to sentence or word-level granularity.
- Future Directions:
- Develop more comprehensive NLP techniques to expand the scope of implicit referencing recognition.
- Explore applicability to other types of digital content (e.g., UI elements, data visualizations).
- Enhance the model's adaptability to semantically and structurally complex texts (e.g., multi-paragraph embedding algorithms like Doc2Vec).
- Limitations:
In summary, this study successfully validated the feasibility of combining speech and gaze modalities, providing a reliable technical foundation for implicit human-computer interaction design and demonstrating potential for broader application scenarios.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can speech and gaze data be combined to achieve implicit human-computer interaction?Category: Social Interaction, Remote Connection, and Relationship ExperienceSimilar questionsarrow_forward
- How can feature extraction between gaze patterns and speech semantics improve implicit annotation precision?Category: Social Interaction, Remote Connection, and Relationship ExperienceSimilar questionsarrow_forward
- Can classifiers combining speech and gaze adapt to complex speech annotation scenarios?Category: Social Interaction, Remote Connection, and Relationship ExperienceSimilar questionsarrow_forward
Practical Problems
1- When quickly browsing text, existing gaze interaction technologies have insufficient efficiency and precision.Category: Social Interaction, Remote Connection, and Relationship ExperienceSimilar questionsarrow_forward
- 100%
An Evaluation of Radar Metaphors for Providing Directional Stimuli Using Non-Verbal Sound
CHI '19· Eye Tracking & Gaze Interaction +1
- 100%
TAGSwipe: Touch Assisted Gaze Swipe for Text Entry
CHI '20· Eye Tracking & Gaze Interaction +1
- 100%
Hummer: Text Entry by Gaze and Hum
CHI '21· Eye Tracking & Gaze Interaction +1
- 100%
EyeSayCorrect: Eye Gaze and Voice Based Hands-free Text Correction for Mobile Devices
IUI '22· Eye Tracking & Gaze Interaction +1
- 67%
M^2Silent: Enabling Multi-user Silent Speech Interactions via Multi-directional Speakers in Shared Spaces
CHI '25· Eye Tracking & Gaze Interaction +2
- 67%
ThumbAir: In-Air Typing for Head Mounted Displays
UbiComp '23· Eye Tracking & Gaze Interaction +2
- 67%
Viago: Exploring Visual-Audio Modality Transitions for Social Media Consumption on the Go
UIST '25· Eye Tracking & Gaze Interaction +2
Based on Jaccard similarity of research subtopics & professions (≥60%)