Satori 悟り: Towards Proactive AR Assistant with Belief-Desire-Intention User Modeling
Authors
Research Background and Issues
- Identified Problems or Challenges:
Current Augmented Reality (AR) assistants are predominantly based on reactive interactions, where the system's behavior is triggered by specific user inputs or environmental changes. This approach faces the following key challenges:
- Lack of proactive interaction, requiring users to frequently issue explicit commands;
- Existing AR assistants often rely on fixed rules, lacking flexibility and scalability;
- Poor performance in continuous guidance for multi-step tasks and multitasking transitions.
- Significance of the Problem: Proactive AR assistants have the potential to provide timely guidance by better understanding users' actions and environments, thereby reducing cognitive load and improving task completion efficiency. The timing of appropriate guidance is particularly critical for user experience and task performance.
- Research Motivation and Related Work: Based on extensive research in virtual assistants and cognitive psychology, the authors propose leveraging the "Belief-Desire-Intention" (BDI) model to enhance the proactivity of AR assistants. While the BDI model has been widely studied in software engineering and cognitive modeling, it has yet to be applied in the AR domain.
Solution
- Proposed Approach: The authors developed a novel AR system, Satori, which integrates the BDI model with large multimodal language models (LLMs), such as GPT-4V, to enable proactive AR interactions.
- Innovations:
- Adapting the BDI model from psychology as a user modeling framework for AR scenarios;
- Enhancing situational, task, and user action insights by combining vision models with LLMs;
- Dynamically generating multimodal content, including text, images, and audio guidance;
- Providing real-time, adaptive task guidance that supports task switching and step-by-step confirmation, improving adaptability to complex environments.
- Implementation Steps:
- BDI Model Adaptation:
- Belief: Deriving the user's current environmental understanding through visual perception modules (scene detection, object recognition) and the user's action history.
- Desire: Using LLMs to infer the user's high-level task goals and refining accuracy through user confirmation.
- Intention: Predicting the user's next actions to provide timely and relevant step-by-step guidance.
- Dynamic Content Generation:
- Utilizing image generation models like DALL-E to create task-specific illustrative images;
- Dynamically configuring the format of guidance (text, audio, tooltips, etc.).
- Timing Prediction and Confirmation Mechanism:
- Implementing real-time action completion detection and prediction to reduce response delays;
- Ensuring guidance accuracy through user confirmation interfaces.
- Interface Design:
- Offering a transparent task tracking and user interaction interface, displaying task goals, pending subtasks, and object locations.
- BDI Model Adaptation:
Research Outcomes
- Specific Results:
- Performance:
- Evaluated using common task datasets (e.g., GTEA) and author-generated video data, Satori achieved high accuracy in task scene understanding, user action prediction, guidance timing, and content generation.
- Compared to a professionally designed Wizard-of-Oz (WoZ) baseline system, Satori demonstrated comparable timing accuracy (81.69%) and an average recall rate of 75.12% for guidance modality prediction in multi-user studies.
- User Studies:
- In a 16-participant user experiment, Satori was rated similarly to the WoZ system in terms of content comprehensibility, task relevance, and response timeliness.
- User satisfaction across multiple dimensions (e.g., guidance timing, usability, and practicality) was comparable to or even better than the baseline; in NASA-TLX cognitive load assessments, Satori showed no significant difference from the WoZ system.
- Text and multimodal content generation received positive feedback from users, with features like arrow-based positioning and illustrative images proving highly valuable in guidance.
- Performance:
- Advantages Over Existing Solutions:
- Does not rely on predefined task rules or manual configurations, offering greater scalability;
- Automates content generation and guidance updates, reducing the need for repetitive design efforts;
- Enhances support for multitasking management and user intention detection, improving responsiveness to real-time contexts.
- Limitations and Future Directions:
- Technical Limitations:
- The system experiences a 2-3 second delay, which may affect guidance for rapid tasks;
- High-level task prediction remains insufficiently accurate in certain scenarios.
- Device Constraints:
- Current AR headsets (e.g., HoloLens 2) have limited fields of view, potentially missing critical user actions;
- More optimized sensor devices (e.g., third-person cameras) are needed to provide comprehensive environmental understanding.
- Future Research Directions:
- Expanding to broader domains (e.g., industrial, medical) to explore domain-specific task applicability;
- Incorporating neural signals or enhanced cross-modal methods to improve the accuracy of user belief modeling;
- Introducing features to support multi-user collaboration and complex social interactions, such as task delegation and group goal modeling.
- Technical Limitations:
The above analysis demonstrates that Satori represents a significant step forward in the design of proactive AR assistance systems, showcasing broad potential and promising applications in terms of generalizability and interaction effectiveness.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can the belief-desire-intention (BDI) model be applied in augmented reality (AR) to improve proactive system interaction?Category: NPC Dialogue and Character Interaction in XRSimilar questionsarrow_forward
- How can combining vision models and large language models (LLMs) improve AR assistants' understanding of user behavior and task contexts?Category: NPC Dialogue and Character Interaction in XRSimilar questionsarrow_forward
- How can multimodal content be dynamically generated to support task switching and continuous guidance for multi-step tasks?Category: NPC Dialogue and Character Interaction in XRSimilar questionsarrow_forward
Practical Problems
1- Users must issue frequent commands, and existing AR assistants struggle to provide continuous guidance for complex tasks.Category: NPC Dialogue and Character Interaction in XRSimilar questionsarrow_forward
- 67%
NavigaTone: Seamlessly Embedding Navigation Cues in Mobile Music Listening
CHI '18· AR Navigation & Context Awareness +1
- 67%
OmniActions: Predicting Digital Actions in Response to Real-World Multimodal Sensory Inputs with LLMs
CHI '24· Human-LLM Collaboration +1
- 67%
Modes of Interaction with Navigation Apps
CHI '25· AR Navigation & Context Awareness +1
- 67%
On the Go with AR: Attention to Virtual and Physical Targets while Varying Augmentation Density
CHI '25· AR Navigation & Context Awareness +1
- 67%
Sensing Noticeability in Ambient Information Environments
CHI '25· AR Navigation & Context Awareness +1
- 67%
Negotiating Dyadic Interactions through the Lens of Augmented Reality Glasses
DIS '23· AR Navigation & Context Awareness +1
- 67%
Walking with Adaptive Augmented Reality Workspaces: Design and Usage Patterns
IUI '19· AR Navigation & Context Awareness +1
- 67%
A First Exploration on the Use of Head-Mounted Augmented Reality in the Context of the Portuguese Military
MobileHCI '23· AR Navigation & Context Awareness +1
- 67%
Context-Aware Online Adaptation of Mixed Reality Interfaces
UIST '19· AR Navigation & Context Awareness +1
Based on Jaccard similarity of research subtopics & professions (≥60%)