ViSTAR: Virtual Skill Training with Augmented Reality with 3D Avatars and LLM coaching agent

Full-Body Interaction & Embodied InputEye Tracking & Gaze InteractionBrain-Computer Interface (BCI) & NeurofeedbackFitness Tracking & Physical Activity MonitoringBehavior Change & Reflection TechnologyAthletes & Fitness EnthusiastsPersonal Trainers & Fitness CoachesAI/ML Researchers & Engineers

Paper Title

ViSTAR: Virtual Skill Training with Augmented Reality with 3D Avatars and LLM coaching agent

Publication Info

  • Topic area: Augmented Reality (AR) and Artificial Intelligence (AI) for motor skill training in sports.
  • Keywords: Augmented Reality, Behavioral Skills Training, 3D motion reconstruction, Large Language Models, basketball training, motor skill acquisition, verbal feedback, visual feedback, self-guided learning, AI coaching.

Background and Problem

  • Problem / challenge: Current sports training systems, particularly in AR/VR, lack personalization, multi-angle visualizations, and actionable feedback for complex motor skills. They often fail to provide nuanced guidance for multi-phase, coordinated movements like advanced basketball techniques.
  • Significance: Effective motor skill training is critical for both recreational and professional athletes, but personalized coaching is resource-intensive and inaccessible to many. AR/AI systems could democratize access to high-quality training.
  • Motivation and related work: Prior AR/VR systems have focused on tactical decision-making or repeatable skills but struggle with ecological feasibility for dynamic sports. AI-driven systems have shown promise in generating feedback, but they often lack the integration of visual and verbal cues or fail to address embodied aspects of learning.

Solution

  • Proposed approach: ViSTAR, an AR-based system leveraging 3D avatars and Large Language Models (LLMs) to provide multi-faceted feedback (visual and verbal) for basketball skill training, grounded in the Behavioral Skills Training (BST) framework.
  • Novelty:
    1. Integration of 3D motion reconstruction with LLM-generated verbal feedback for actionable coaching.
    2. A feedback design framework combining holistic (flow/rhythm) and localized (joint-level) guidance.
    3. A novel method for generating verbal feedback by linking joint-level motion analysis with LLMs, using Random Forest classifiers for prioritization.
    4. Empirical evaluation of AI vs. coach feedback and AR-guided training through user studies.
  • Procedure and key techniques:
    • Instruction: 3D avatars demonstrate ideal motions with multi-angle views and motion trails.
    • Modeling: Segmented expert overlays and pose-matching interfaces provide step-by-step guidance.
    • Rehearsal: Users record and review their own motions with synchronized comparisons to expert models.
    • Feedback: Multi-faceted feedback includes visual heatmaps, timing judgments, and LLM-generated verbal coaching.

Results

  • Concrete findings:
    • ViSTAR reduced angular errors in user motions compared to self-observation (Trial 1: 10.80° vs. 14.95°, p = 0.0155; Trial 2: 11.32° vs. 14.60°, p = 0.0131).
    • AI-generated feedback was preferred over real coach feedback in 41.2% of cases, citing clarity and actionability.
    • 100% of participants reported ViSTAR helped them recognize errors, and 94% found it useful for corrections.
  • Advantage over baselines:
    • ViSTAR outperformed traditional self-observation in error recognition and correction.
    • AI feedback was rated higher than coach feedback in clarity, identifiability, and actionability.
  • Experiments / evaluation:
    • Two user studies with 16 basketball players (2-18 years of experience).
    • Tasks included verbal feedback comparison (AI vs. coach) and skill training (self-observation vs. ViSTAR).
    • Metrics: angular deviation, usability, engagement, and feedback quality.
  • Limitations and future work:
    • Small sample size and short-term evaluation limit generalizability.
    • Current system lacks multi-player scenarios, ball tracking, and long-term training validation.
    • Visual clutter and pacing issues in dynamic motions need refinement.
    • Future directions include integrating multimodal inputs (e.g., voice commands) and adaptive feedback strategies.

Summary

ViSTAR is an AR-based skill training system that combines 3D motion reconstruction with LLM-powered verbal feedback to support self-guided basketball practice. Grounded in the Behavioral Skills Training framework, it provides multi-faceted feedback to help users recognize and correct errors in posture, timing, and coordination. User studies demonstrated that ViSTAR improves motion accuracy and is preferred over traditional self-observation and coach feedback for its clarity and actionability. While exploratory findings highlight its promise, future work is needed to address ecological validity, scalability, and dynamic motion challenges. ViSTAR represents a step toward democratizing access to high-quality, embodied skill training through AR and AI integration.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222543/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790634
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Full-Body Interaction & Embodied Input, Eye Tracking & Gaze Interaction, Brain-Computer Interface (BCI) & Neurofeedback, Fitness Tracking & Physical Activity Monitoring
work
Professions
Athletes & Fitness Enthusiasts, Personal Trainers & Fitness Coaches, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers