GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented Reality
Authors
Title of the Paper
GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented Reality
Paper Information
- Research Area: Augmented reality technology, natural language processing, voice assistants, human-computer interaction
- Keywords: augmented reality, voice assistant, multimodal input, gaze tracking, pointing gesture recognition, large language models
Research Background and Problem
- Current mainstream voice assistants (e.g., Siri and Alexa) lack an understanding of users' spatial and temporal contexts, resulting in limited performance when responding to ambiguous queries, such as those involving pronouns.
- Pronouns are a critical component of human daily communication, yet existing voice assistants provide weak support for them and fail to effectively resolve unclear pronoun references (i.e., pronoun disambiguation).
- The significance of this issue lies in enhancing voice assistants' ability to naturally understand context and interact like humans, thereby improving user experience.
Solution
Methodology and Innovations
- Introduction of GazePointAR:
- GazePointAR is a context-aware voice assistant designed for wearable augmented reality, combining gaze tracking, pointing gestures, and conversational history to address ambiguous pronoun references in user queries.
- By leveraging real-time computer vision technology, it identifies objects in the user's field of view and utilizes large language models (GPT-3 or GPT-3.5) to generate responses.
- A novel multimodal interaction method is proposed, enabling users to interact with the system through gaze, pointing, and brief natural language queries.
Implementation Steps and Key Technologies
- User Activation: The system is activated by the user saying "Hey Glass," entering a "listening mode" to detect voice queries.
- Scene Analysis:
- Captures images of the user's field of view.
- Applies computer vision techniques to extract objects, text, and facial information from the images.
- Gaze Tracking and Gesture Recognition:
- Uses built-in modules to obtain the 3D coordinates of the user's gaze position.
- Implements pointing gesture recognition and combines it with real-time spatial mesh detection to identify the user's pointing target.
- Query Construction and Pronoun Replacement:
- Generates clear phrases to replace ambiguous pronouns based on gaze and gesture information.
- Enhances pronoun disambiguation intelligence using prompt engineering and GPT technology.
- Answer Generation and Feedback:
- Sends the optimized query to the large language model (GPT-3.5) to retrieve an answer, which is then read aloud to the user via text-to-speech functionality.
Research Outcomes
Specific Results
- A fully functional prototype of the context-aware voice assistant was developed and tested through two studies:
- Laboratory Study: Compared the performance of GazePointAR with existing voice assistants (Google Voice Assistant and Google Lens).
- First-Person Diary Study: Documented researchers' experiences using GazePointAR in real-world scenarios over five consecutive days.
Experimental and Evaluation Results
-
Laboratory Study Results:
- Compared to Google Voice Assistant and Google Lens, users found GazePointAR to be more natural and human-like.
- GazePointAR excelled in handling pronoun-based voice queries, such as quickly and accurately identifying objects referred to by pronouns like "this" or "that."
- Key issues included the singularity of answers and limited ability to access historical context.
-
First-Person Diary Study Results:
- Performed well in real-world scenarios involving complex queries (e.g., "This dress is too expensive; recommend similar brands"), but showed limitations in tasks lacking contextual information or requiring multiple pronoun references.
- Participants were satisfied with its natural and intuitive interaction capabilities but noted the need for improvements in continuous gaze tracking and refined target identification.
Advantages Over Existing Solutions
- Multimodal Support: Integrates gaze tracking and pointing gestures, enabling context-awareness beyond traditional voice assistants.
- Enhanced Pronoun Recognition: Improves users' ability to construct simple queries, allowing interaction with devices through ambiguous expressions in natural language.
- Real-Time Responsiveness and Interactivity: Provides rapid responses and the ability to explain answers.
Limitations and Future Directions
- Limitations:
- Difficulty handling queries involving multiple pronouns and historical references.
- Insufficient recognition capabilities for complex objects.
- Potential privacy and comfort concerns when used in public settings.
- Extended gaze fixation on objects may lead to user fatigue.
- Future Directions:
- Support continuous gaze tracking to capture dynamic contexts.
- Offer multiple query answers and allow users to select or edit responses.
- Optimize machine learning models to recognize a broader range of object categories.
- Balance privacy protection with functionality expansion.
Conclusion
This paper presents a context-aware voice assistant based on augmented reality devices, capable of resolving pronoun disambiguation through multimodal input. The study demonstrates that GazePointAR enhances the naturalness and accuracy of voice interactions and explores the innovative potential of combining augmented reality technology with large language models. With further technological optimization, future context-aware voice assistants will provide broader support for natural language queries and dynamic interactions, significantly improving user experience.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- In wearable AR, how can multimodal interaction (gaze, pointing gestures, and dialogue history) resolve pronoun ambiguity?Category: XR Input, Tracking, and Spatial InteractionSimilar questionsarrow_forward
- Compared with existing voice assistants, can multimodal interaction significantly improve pronoun reference accuracy?Category: XR Input, Tracking, and Spatial InteractionSimilar questionsarrow_forward
- How do users evaluate the naturalness and interaction capability of GazePointAR in real-world scenarios?Category: XR Input, Tracking, and Spatial InteractionSimilar questionsarrow_forward
Practical Problems
1- Pronoun ambiguity when interacting with voice assistants often leads to misunderstanding or inefficient communication.Category: XR Input, Tracking, and Spatial InteractionSimilar questionsarrow_forward
- 80%
Pinpointing: Precise Head- and Eye-Based Target Selection for Augmented Reality
CHI '18· Eye Tracking & Gaze Interaction +1
- 80%
Stretch Gaze Targets Out: Experimenting with Target Sizes for Gaze-Enabled Interfaces on Mobile Devices
CHI '25· Eye Tracking & Gaze Interaction +1
Based on Jaccard similarity of research subtopics & professions (≥60%)