CARING-AI: Towards Authoring Context-aware Augmented Reality INstruction through Generative Artificial Intelligence
Authors
Research Background and Issues
-
Issues and Challenges:
- Augmented Reality (AR) instructions offer an immersive way of learning and interaction, helping users understand complex tasks by rendering digital content into real-world environments. However, current methods utilizing AI-generated content (AIGC) lack sufficient contextual awareness in AR applications.
- Existing approaches, including motion-capture-based demonstration programming, can generate AR animation content but are limited by hardware complexity, contextual applicability, and the need for specialized expertise.
- Generative AI-based methods lack spatial and temporal alignment between the generated instructions and the user's actual environment.
-
Significance:
- Generating context-aware instructions tailored to users and their environments in real-world scenarios is crucial for effective guidance.
- Context awareness can significantly enhance the adaptability, learning outcomes, and interactivity of AR instructions.
-
Research Motivation and Related Work:
- Motivation: To leverage generative AI (e.g., text generation models and motion diffusion models) to produce AR instructions better suited to complex real-world scenarios, eliminating the need for complex hardware and programming skills.
- Related Work: Earlier explorations using Generative Adversarial Networks (GANs) to generate content for AR applications primarily focused on static 3D images or simple environmental setups. CARING-AI aims to achieve dynamic, context-aware 3D humanoid animations.
Solution
-
Methods and Solutions:
- The CARING-AI system proposes a method for generating context-aware humanoid AR animation instructions using generative AI (e.g., ChatGPT and diffusion models).
-
Innovations:
- Developed a workflow that eliminates the need for programming and motion capture, generating AR instructions adaptable to human, environmental, and system contexts.
- Implemented a dynamic smoothing algorithm based on diffusion models, ensuring natural temporal transitions in multi-step animations.
- Provided a user-friendly AR interface supporting task text input, environment scanning, instruction animation generation, and visualization/editing of instructions.
-
Implementation Steps:
- Generating Text Instructions:
- ChatGPT generates step-by-step task instructions, which users can edit and adjust.
- Environment Scanning:
- Users scan the real environment using HoloLens2 to provide object location information.
- Animation Generation:
- Based on user-provided text instructions and context, the diffusion model generates humanoid motion animations, including local and global spatial and temporal context mapping.
- Visualization and Editing:
- An interactive AR interface allows users to adjust, modify, and preview animations.
- Generating Text Instructions:
Research Outcomes
-
Specific Achievements:
- Defined a design space for addressing diverse contexts and content scopes.
- Successfully developed a system enabling users to generate and edit AR instruction animations, including local (e.g., hand-object interaction) and global (e.g., path navigation in environments) contextual adaptation.
-
Advantages Over Existing Solutions:
- Compared to programming-based or demonstration-based content generation methods, CARING-AI does not require users to have specialized knowledge or expensive motion capture equipment.
- Supports more complex and natural multi-step task animations, addressing deficiencies in temporal continuity and smoothness in existing methods.
-
Experimental Results:
- Experiment 1: Algorithm Performance Evaluation: Compared to baseline algorithms (GMD), CARING-AI significantly improved animation temporal continuity (transition distance reduced to 0.03m) and achieved high spatial alignment accuracy (error less than 0.1m).
- Experiment 2: Usability Validation: Users rated CARING-AI's system usability at 83.21 (out of 100), with high recognition for animation quality and generation convenience.
- Experiment 3: Interaction Testing: Compared to demonstration-based programming, CARING-AI significantly reduced task completion time and error rates, with user feedback indicating lower psychological and physical burdens than baseline methods.
-
Limitations and Future Directions:
- Limitations:
- The system's current ability to generate hand-object interaction is limited, unable to support complex object deformation and deep interactions.
- Certain hardware constraints (e.g., device field of view, computational latency) may impact user experience.
- Future Directions:
- Enhance high-fidelity hand-object interaction capabilities by introducing algorithms supporting flexible and segmented objects.
- Integrate more types of generative content (e.g., images, audio, video) to enrich AR scenarios.
- Improve real-time system performance and explore possibilities of combining larger models and cloud computing.
- Limitations:
This paper demonstrates the potential of generative AI in AR instruction generation through the CARING-AI system, offering new insights and application directions for future AR development based on generative algorithms.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can generative AI produce AR task instruction animations adapted to complex real-world scenarios?Category: AR Task Guidance, Tutorials, and Prompt DesignSimilar questionsarrow_forward
- How can generated AR instructions dynamically adapt to diverse user, environment, and system contexts?Category: AR Task Guidance, Tutorials, and Prompt DesignSimilar questionsarrow_forward
- What algorithms can improve spatiotemporal continuity and naturalness of AR animations?Category: AR Task Guidance, Tutorials, and Prompt DesignSimilar questionsarrow_forward
Practical Problems
1- Users struggle to generate context-adaptive AR guidance content in real environments.Category: AR Task Guidance, Tutorials, and Prompt DesignSimilar questionsarrow_forward
- 83%
MUD: Towards a Large-Scale and Noise-Filtered UI Dataset for Modern Style UI Modeling
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 83%
How the Role of Generative AI Shapes Perceptions of Value in Human-AI Collaborative Work
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 83%
agentAR: Creating Augmented Reality Applications with Tool-Augmented LLM-based Autonomous Agents
UIST '25· AR Navigation & Context Awareness +2
- 71%
ReactGenie: A Development Framework for Complex Multimodal Interactions Using Large Language Models
CHI '24· Voice User Interface (VUI) Design +2
- 71%
AI-Augmented Brainwriting: Investigating the use of LLMs in group ideation
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 71%
Generative AI in Knowledge Work: Design Implications for Data Navigation and Decision-Making
CHI '25· Generative AI (Text, Image, Music, Video) +2
- 71%
Beyond Code Generation: LLM-supported Exploration of the Program Design Space
CHI '25· Generative AI (Text, Image, Music, Video) +2
- 71%
Prototyping with Prompts: Emerging Approaches and Challenges in Generative AI Design for Collaborative Software Teams
CHI '25· Generative AI (Text, Image, Music, Video) +2
- 71%
Creating Design Resources to Scaffold the Ideation of AI Concepts
DIS '23· Generative AI (Text, Image, Music, Video) +2
- 71%
GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and Learning
UIST '24· Generative AI (Text, Image, Music, Video) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)