ARify: Leveraging Narrated Instructional Videos to Create Augmented Reality Tutorials for Procedural Tasks
Authors
Paper Title
ARify: Leveraging Narrated Instructional Videos to Create Augmented Reality Tutorials for Procedural Tasks
Publication Info
- Topic area: Augmented Reality tutorial creation from instructional videos
- Keywords: Augmented Reality, instructional videos, procedural tasks, tutorial authoring, vision-language models, AR representations, usability, automation, design space
Background and Problem
- Problem / challenge: Creating Augmented Reality (AR) tutorials for procedural tasks requires significant AR authoring expertise, which limits accessibility for domain experts without technical skills.
- Significance: AR tutorials improve task learning by providing situated, step-by-step guidance, reducing cognitive load, and enhancing learning outcomes. Lowering the barrier to AR tutorial creation can democratize access to this effective learning medium.
- Motivation and related work: Prior systems either rely on manual AR authoring in game engines or use rule-based methods that are task-specific and lack generalizability. While demonstration-based and narrative-enhanced systems improve interpretability, they still require AR expertise. Emerging vision-language models (VLMs) show promise for automating tutorial generation but have yet to be applied to narrated instructional videos for AR tutorial creation.
Solution
- Proposed approach: ARify, a semi-automatic authoring system that transforms narrated instructional videos into AR tutorials using a vision-language model (VLM) and interactive editing tools.
- Novelty:
- A design space linking instructional intents, tactics, and AR representations derived from content analysis of instructional videos.
- A VLM-based system that automatically generates tutorial structures and AR representations from narrated instructional videos.
- Interactive interfaces (Tutorial Structure Editor and AR Editor) for refining and customizing AR tutorials.
- Empirical validation through numerical and user studies demonstrating effectiveness and usability.
- Procedure and key techniques:
- Content Analysis: Identify instructional intents and tactics in videos and map them to AR representations.
- Tutorial Planner: Use a VLM to parse videos into hierarchical task → step → action structures and recommend AR representations.
- AR Builder: Configure AR representations using perception modules for object detection, pose estimation, and spatial placement.
- Authoring Interfaces: Provide tools for users to refine tutorial structures and AR elements.
- Deployment: Export AR tutorials for head-mounted displays (HMDs) with spatially grounded, stepwise guidance.
Results
- Concrete findings:
- ARify achieved sequence accuracy of 88.8% and F1 scores of 90.1% on average across three task types (assembly, repair, training).
- Strong performance across instructional tactics, with F1 scores above 80% for all categories.
- Usability study participants rated the system highly (SUS score: 80.00 ± 14.85) and found it intuitive and effective for creating AR tutorials.
- Advantage over baselines:
- Few-shot demonstrations significantly improved performance compared to zero-shot conditions (all p < 0.001).
- ARify’s multi-task demonstration strategy outperformed task-specific configurations in generalizability across task types (all p < 0.01).
- Experiments / evaluation:
- Numerical study: Evaluated ARify on 36 additional instructional videos (2371 annotated tactics) for planning accuracy and tactic-level performance.
- Usability study: Conducted with 18 participants creating AR tutorials for three machine tasks, assessing ease of use, satisfaction, and effectiveness.
- Limitations and future work:
- Challenges in handling dense clutter, small objects, and subtle motor skills due to limitations in perception modules and VLM segmentation.
- Limited AR representation library; future work could expand to include more specialized elements.
- Current implementation targets AR-HMDs; future extensions could support mobile AR platforms and multimodal feedback.
Summary
ARify is a semi-automatic system that transforms narrated instructional videos into AR tutorials, leveraging a vision-language model and interactive authoring tools. It demonstrated strong performance in generating accurate and usable AR tutorials across diverse procedural tasks, significantly lowering the entry barrier for novice users. While current limitations include challenges with complex scenes and a constrained AR representation library, ARify provides a robust foundation for scalable AR tutorial creation and offers opportunities for future enhancements in generalizability, representation diversity, and cross-platform adaptability.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
Spatialstrates: Cross-Reality Collaboration through Spatial Hypermedia
UIST '25· Social & Collaborative VR +2
- 67%
RemapVR: An Immersive Authoring Tool for Rapid Prototyping of Remapped Interaction in VR
CHI '25· Mixed Reality Workspaces +1
- 67%
Traversing Dual Realities: Investigating Techniques for Transitioning 3D Objects between Desktop and Augmented Reality Environments
CHI '25· AR Navigation & Context Awareness +1
- 67%
Beyond the Input Stream: Making Text Entry Evaluations More Flexible with Transcription Sequences
UIST '19· Prototyping & User Testing
- 67%
DepthLab: Real-time 3D Interaction with Depth Maps for Mobile Augmented Reality
UIST '20· AR Navigation & Context Awareness +1
- 67%
Rapido: Prototyping Interactive AR Experiences through Programming by Demonstration
UIST '21· AR Navigation & Context Awareness +1
- 67%
AUIT – the Adaptive User Interfaces Toolkit for Designing XR Applications
UIST '22· AR Navigation & Context Awareness +1
Based on Jaccard similarity of research subtopics & professions (≥60%)