Guided Reality: Generating Visually-Enriched AR Task Guidance with LLMs and Vision Models

AR Navigation & Context AwarenessHuman-LLM CollaborationVocational Trainers & CoachesSoftware Engineers & Developers

Large language models (LLMs) have enabled the automatic generation of step-by-step augmented reality (AR) instructions for a wide range of physical tasks. However, existing LLM-based AR guidance often lacks rich visual augmentations to effectively embed instructions into spatial context for a better user understanding. We present Guided Reality, a fully automated AR system that generates embedded and dynamic visual guidance based on step-by-step instructions. Our system integrates LLMs and vision models to: 1) generate multi-step instructions from user queries, 2) identify appropriate types of visual guidance, 3) extract spatial information about key interaction points in the real world, and 4) embed visual guidance in physical space to support task execution. Drawing from a corpus of user manuals, we define five categories of visual guidance and propose an identification strategy based on the current step. We evaluate the system through a user study (N=16), completing real-world tasks and exploring the system in the wild. Additionally, four instructors shared insights on how Guided Reality could be integrated into their training workflows.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/uist/206890/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3746059.3747784
At a Glance

Paper Snapshot

fact_check
dataset
Source
UIST
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
AR Navigation & Context Awareness, Human-LLM Collaboration
work
Professions
Vocational Trainers & Coaches, Software Engineers & Developers
article
Content Status
Abstract only
hub
Related Papers
0 related papers