CARING-AI: Towards Authoring Context-aware Augmented Reality INstruction through Generative Artificial Intelligence

AR Navigation & Context AwarenessGenerative AI (Text, Image, Music, Video)Human-LLM CollaborationSoftware Engineers & DevelopersUI/UX DesignersAI/ML Researchers & Engineers

Research Background and Issues

  • Issues and Challenges:

    • Augmented Reality (AR) instructions offer an immersive way of learning and interaction, helping users understand complex tasks by rendering digital content into real-world environments. However, current methods utilizing AI-generated content (AIGC) lack sufficient contextual awareness in AR applications.
    • Existing approaches, including motion-capture-based demonstration programming, can generate AR animation content but are limited by hardware complexity, contextual applicability, and the need for specialized expertise.
    • Generative AI-based methods lack spatial and temporal alignment between the generated instructions and the user's actual environment.
  • Significance:

    • Generating context-aware instructions tailored to users and their environments in real-world scenarios is crucial for effective guidance.
    • Context awareness can significantly enhance the adaptability, learning outcomes, and interactivity of AR instructions.
  • Research Motivation and Related Work:

    • Motivation: To leverage generative AI (e.g., text generation models and motion diffusion models) to produce AR instructions better suited to complex real-world scenarios, eliminating the need for complex hardware and programming skills.
    • Related Work: Earlier explorations using Generative Adversarial Networks (GANs) to generate content for AR applications primarily focused on static 3D images or simple environmental setups. CARING-AI aims to achieve dynamic, context-aware 3D humanoid animations.

Solution

  • Methods and Solutions:

    • The CARING-AI system proposes a method for generating context-aware humanoid AR animation instructions using generative AI (e.g., ChatGPT and diffusion models).
  • Innovations:

    • Developed a workflow that eliminates the need for programming and motion capture, generating AR instructions adaptable to human, environmental, and system contexts.
    • Implemented a dynamic smoothing algorithm based on diffusion models, ensuring natural temporal transitions in multi-step animations.
    • Provided a user-friendly AR interface supporting task text input, environment scanning, instruction animation generation, and visualization/editing of instructions.
  • Implementation Steps:

    1. Generating Text Instructions:
      • ChatGPT generates step-by-step task instructions, which users can edit and adjust.
    2. Environment Scanning:
      • Users scan the real environment using HoloLens2 to provide object location information.
    3. Animation Generation:
      • Based on user-provided text instructions and context, the diffusion model generates humanoid motion animations, including local and global spatial and temporal context mapping.
    4. Visualization and Editing:
      • An interactive AR interface allows users to adjust, modify, and preview animations.

Research Outcomes

  • Specific Achievements:

    • Defined a design space for addressing diverse contexts and content scopes.
    • Successfully developed a system enabling users to generate and edit AR instruction animations, including local (e.g., hand-object interaction) and global (e.g., path navigation in environments) contextual adaptation.
  • Advantages Over Existing Solutions:

    • Compared to programming-based or demonstration-based content generation methods, CARING-AI does not require users to have specialized knowledge or expensive motion capture equipment.
    • Supports more complex and natural multi-step task animations, addressing deficiencies in temporal continuity and smoothness in existing methods.
  • Experimental Results:

    • Experiment 1: Algorithm Performance Evaluation: Compared to baseline algorithms (GMD), CARING-AI significantly improved animation temporal continuity (transition distance reduced to 0.03m) and achieved high spatial alignment accuracy (error less than 0.1m).
    • Experiment 2: Usability Validation: Users rated CARING-AI's system usability at 83.21 (out of 100), with high recognition for animation quality and generation convenience.
    • Experiment 3: Interaction Testing: Compared to demonstration-based programming, CARING-AI significantly reduced task completion time and error rates, with user feedback indicating lower psychological and physical burdens than baseline methods.
  • Limitations and Future Directions:

    • Limitations:
      • The system's current ability to generate hand-object interaction is limited, unable to support complex object deformation and deep interactions.
      • Certain hardware constraints (e.g., device field of view, computational latency) may impact user experience.
    • Future Directions:
      • Enhance high-fidelity hand-object interaction capabilities by introducing algorithms supporting flexible and segmented objects.
      • Integrate more types of generative content (e.g., images, audio, video) to enrich AR scenarios.
      • Improve real-time system performance and explore possibilities of combining larger models and cloud computing.

This paper demonstrates the potential of generative AI in AR instruction generation through the CARING-AI system, offering new insights and application directions for future AR development based on generative algorithms.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189096/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713348
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
AR Navigation & Context Awareness, Generative AI (Text, Image, Music, Video), Human-LLM Collaboration
work
Professions
Software Engineers & Developers, UI/UX Designers, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers