ViSTAR: Virtual Skill Training with Augmented Reality with 3D Avatars and LLM coaching agent
Authors
Paper Title
ViSTAR: Virtual Skill Training with Augmented Reality with 3D Avatars and LLM coaching agent
Publication Info
- Topic area: Augmented Reality (AR) and Artificial Intelligence (AI) for motor skill training in sports.
- Keywords: Augmented Reality, Behavioral Skills Training, 3D motion reconstruction, Large Language Models, basketball training, motor skill acquisition, verbal feedback, visual feedback, self-guided learning, AI coaching.
Background and Problem
- Problem / challenge: Current sports training systems, particularly in AR/VR, lack personalization, multi-angle visualizations, and actionable feedback for complex motor skills. They often fail to provide nuanced guidance for multi-phase, coordinated movements like advanced basketball techniques.
- Significance: Effective motor skill training is critical for both recreational and professional athletes, but personalized coaching is resource-intensive and inaccessible to many. AR/AI systems could democratize access to high-quality training.
- Motivation and related work: Prior AR/VR systems have focused on tactical decision-making or repeatable skills but struggle with ecological feasibility for dynamic sports. AI-driven systems have shown promise in generating feedback, but they often lack the integration of visual and verbal cues or fail to address embodied aspects of learning.
Solution
- Proposed approach: ViSTAR, an AR-based system leveraging 3D avatars and Large Language Models (LLMs) to provide multi-faceted feedback (visual and verbal) for basketball skill training, grounded in the Behavioral Skills Training (BST) framework.
- Novelty:
- Integration of 3D motion reconstruction with LLM-generated verbal feedback for actionable coaching.
- A feedback design framework combining holistic (flow/rhythm) and localized (joint-level) guidance.
- A novel method for generating verbal feedback by linking joint-level motion analysis with LLMs, using Random Forest classifiers for prioritization.
- Empirical evaluation of AI vs. coach feedback and AR-guided training through user studies.
- Procedure and key techniques:
- Instruction: 3D avatars demonstrate ideal motions with multi-angle views and motion trails.
- Modeling: Segmented expert overlays and pose-matching interfaces provide step-by-step guidance.
- Rehearsal: Users record and review their own motions with synchronized comparisons to expert models.
- Feedback: Multi-faceted feedback includes visual heatmaps, timing judgments, and LLM-generated verbal coaching.
Results
- Concrete findings:
- ViSTAR reduced angular errors in user motions compared to self-observation (Trial 1: 10.80° vs. 14.95°, p = 0.0155; Trial 2: 11.32° vs. 14.60°, p = 0.0131).
- AI-generated feedback was preferred over real coach feedback in 41.2% of cases, citing clarity and actionability.
- 100% of participants reported ViSTAR helped them recognize errors, and 94% found it useful for corrections.
- Advantage over baselines:
- ViSTAR outperformed traditional self-observation in error recognition and correction.
- AI feedback was rated higher than coach feedback in clarity, identifiability, and actionability.
- Experiments / evaluation:
- Two user studies with 16 basketball players (2-18 years of experience).
- Tasks included verbal feedback comparison (AI vs. coach) and skill training (self-observation vs. ViSTAR).
- Metrics: angular deviation, usability, engagement, and feedback quality.
- Limitations and future work:
- Small sample size and short-term evaluation limit generalizability.
- Current system lacks multi-player scenarios, ball tracking, and long-term training validation.
- Visual clutter and pacing issues in dynamic motions need refinement.
- Future directions include integrating multimodal inputs (e.g., voice commands) and adaptive feedback strategies.
Summary
ViSTAR is an AR-based skill training system that combines 3D motion reconstruction with LLM-powered verbal feedback to support self-guided basketball practice. Grounded in the Behavioral Skills Training framework, it provides multi-faceted feedback to help users recognize and correct errors in posture, timing, and coordination. User studies demonstrated that ViSTAR improves motion accuracy and is preferred over traditional self-observation and coach feedback for its clarity and actionability. While exploratory findings highlight its promise, future work is needed to address ecological validity, scalability, and dynamic motion challenges. ViSTAR represents a step toward democratizing access to high-quality, embodied skill training through AR and AI integration.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)