ARify: Leveraging Narrated Instructional Videos to Create Augmented Reality Tutorials for Procedural Tasks

AR Navigation & Context AwarenessPrototyping & User TestingMixed Reality WorkspacesUniversity Professors & ResearchersSoftware Engineers & DevelopersUI/UX Designers

Paper Title

ARify: Leveraging Narrated Instructional Videos to Create Augmented Reality Tutorials for Procedural Tasks

Publication Info

  • Topic area: Augmented Reality tutorial creation from instructional videos
  • Keywords: Augmented Reality, instructional videos, procedural tasks, tutorial authoring, vision-language models, AR representations, usability, automation, design space

Background and Problem

  • Problem / challenge: Creating Augmented Reality (AR) tutorials for procedural tasks requires significant AR authoring expertise, which limits accessibility for domain experts without technical skills.
  • Significance: AR tutorials improve task learning by providing situated, step-by-step guidance, reducing cognitive load, and enhancing learning outcomes. Lowering the barrier to AR tutorial creation can democratize access to this effective learning medium.
  • Motivation and related work: Prior systems either rely on manual AR authoring in game engines or use rule-based methods that are task-specific and lack generalizability. While demonstration-based and narrative-enhanced systems improve interpretability, they still require AR expertise. Emerging vision-language models (VLMs) show promise for automating tutorial generation but have yet to be applied to narrated instructional videos for AR tutorial creation.

Solution

  • Proposed approach: ARify, a semi-automatic authoring system that transforms narrated instructional videos into AR tutorials using a vision-language model (VLM) and interactive editing tools.
  • Novelty:
    1. A design space linking instructional intents, tactics, and AR representations derived from content analysis of instructional videos.
    2. A VLM-based system that automatically generates tutorial structures and AR representations from narrated instructional videos.
    3. Interactive interfaces (Tutorial Structure Editor and AR Editor) for refining and customizing AR tutorials.
    4. Empirical validation through numerical and user studies demonstrating effectiveness and usability.
  • Procedure and key techniques:
    1. Content Analysis: Identify instructional intents and tactics in videos and map them to AR representations.
    2. Tutorial Planner: Use a VLM to parse videos into hierarchical task → step → action structures and recommend AR representations.
    3. AR Builder: Configure AR representations using perception modules for object detection, pose estimation, and spatial placement.
    4. Authoring Interfaces: Provide tools for users to refine tutorial structures and AR elements.
    5. Deployment: Export AR tutorials for head-mounted displays (HMDs) with spatially grounded, stepwise guidance.

Results

  • Concrete findings:
    • ARify achieved sequence accuracy of 88.8% and F1 scores of 90.1% on average across three task types (assembly, repair, training).
    • Strong performance across instructional tactics, with F1 scores above 80% for all categories.
    • Usability study participants rated the system highly (SUS score: 80.00 ± 14.85) and found it intuitive and effective for creating AR tutorials.
  • Advantage over baselines:
    • Few-shot demonstrations significantly improved performance compared to zero-shot conditions (all p < 0.001).
    • ARify’s multi-task demonstration strategy outperformed task-specific configurations in generalizability across task types (all p < 0.01).
  • Experiments / evaluation:
    • Numerical study: Evaluated ARify on 36 additional instructional videos (2371 annotated tactics) for planning accuracy and tactic-level performance.
    • Usability study: Conducted with 18 participants creating AR tutorials for three machine tasks, assessing ease of use, satisfaction, and effectiveness.
  • Limitations and future work:
    • Challenges in handling dense clutter, small objects, and subtle motor skills due to limitations in perception modules and VLM segmentation.
    • Limited AR representation library; future work could expand to include more specialized elements.
    • Current implementation targets AR-HMDs; future extensions could support mobile AR platforms and multimodal feedback.

Summary

ARify is a semi-automatic system that transforms narrated instructional videos into AR tutorials, leveraging a vision-language model and interactive authoring tools. It demonstrated strong performance in generating accurate and usable AR tutorials across diverse procedural tasks, significantly lowering the entry barrier for novice users. While current limitations include challenges with complex scenes and a constrained AR representation library, ARify provides a robust foundation for scalable AR tutorial creation and offers opportunities for future enhancements in generalizability, representation diversity, and cross-platform adaptability.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222666/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790715
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
AR Navigation & Context Awareness, Prototyping & User Testing, Mixed Reality Workspaces
work
Professions
University Professors & Researchers, Software Engineers & Developers, UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
7 related papers