Toward Equitable ASL Education: Egocentric Stereo Sensing with LLM Feedback for Error-Aware Learning
Authors
Paper Title
Toward Scalable ASL Education: Egocentric Stereo Sensing with LLM Feedback for Error-Aware Learning
Publication Info
- Topic area: AI-driven tools for American Sign Language (ASL) education.
- Keywords: ASL learning, egocentric stereo sensing, error detection, large language models, feedback generation, sign language education, human-computer interaction, embodied learning, AI pedagogy, accessibility.
Background and Problem
- Problem / challenge: Existing ASL learning tools focus on recognition rather than providing structured, corrective feedback. Learners lack timely, individualized feedback, especially for manual parameters like handshape, orientation, location, and movement.
- Significance: ASL is vital for communication and cultural identity in the Deaf community, yet structural barriers in education limit access to effective learning resources. Scalable solutions are needed to address the feedback gap for over 500,000 learners in the U.S.
- Motivation and related work: Prior systems rely on front-view cameras or wearables, which are either cumbersome or fail to capture 3D spatial nuances. Feedback mechanisms are often qualitative or binary, lacking pedagogical grounding. This paper builds on SLA theory and ASL curricula to design a system that provides parameter-specific, actionable feedback.
Solution
- Proposed approach: An egocentric stereo camera-based ASL learning system that integrates 3D hand reconstruction, multi-task error detection, and LLM-driven feedback generation.
- Novelty:
- Introduction of the first expert-supervised feedback library tailored for ASL training guidance.
- Development of an end-to-end egocentric stereo ASL learning system combining sensing, error detection, and structured LLM prompting.
- Empirical grounding of design goals through formative studies with ASL instructors and learners.
- Demonstration of robust technical and pedagogical behavior in initial feasibility evaluations.
- Procedure and key techniques:
- Use of a head-mounted stereo camera for egocentric 3D hand reconstruction.
- Multi-task error detection across manual parameters (handshape, orientation, location, movement) using transformer-based modeling.
- Structured feedback generation via GPT-4, conditioned on parameter-specific error signals.
- Confidence gating to suppress uncertain detections and reduce hallucinations in feedback.
Results
- Concrete findings:
- Video-level success rate (V-SR) of 79.14% with a false positive rate (FPR) of 5.93%.
- Feedback coverage of 80.10% for instructor-marked errors and a hallucination rate of 5.63%.
- Scoring-level metrics show high stability, with 90.1% accuracy for orientation errors ≤20° and 89.2% accuracy for location errors ≤10 cm.
- Advantage over baselines:
- Substantially reduced hallucinations compared to 2D mono-view, front-view, rule-based, and zero-shot GPT baselines.
- Higher recall and lower false positive rates across all manual parameters.
- Experiments / evaluation:
- Two studies: formative needs assessment with 45 participants (15 instructors, 30 learners) and system evaluation with 13 Deaf participants practicing 230 signs.
- Metrics include scoring-level stability, video-level agreement, and feedback-level coverage/hallucination.
- Ablations highlight the importance of stereo sensing, multi-task detection heads, and confidence gating.
- Limitations and future work:
- Current evaluation is short-term and limited to manual parameters; non-manual markers like facial expressions are not integrated.
- Small sample size and lack of longitudinal studies limit generalizability.
- Future work includes multi-view sensing for facial cues, larger-scale deployments, and alignment with diverse ASL curricula.
Summary
This paper introduces an egocentric stereo ASL learning system that provides structured, parameter-specific feedback using AI-driven techniques. The system addresses gaps in ASL education by combining 3D sensing, error detection, and LLM-based feedback generation. Evaluation with Deaf participants demonstrates robust technical behavior, high feedback reliability, and alignment with instructor judgments. While limited to manual parameters and short-term studies, the system establishes a foundation for scalable, culturally grounded ASL education and offers transferable design principles for AI-powered embodied skill learning.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)