Giving Meaning to Movements: Challenges and Opportunities in Expanding Communication by Pairing Unaided AAC with Speech Generated Messages
Authors
Paper Title
Giving Meaning to Movements: Challenges and Opportunities in Expanding Communication by Pairing Unaided AAC with Speech Generated Messages
Publication Info
- Topic area: Combining aided and unaided augmentative and alternative communication (AAC) for individuals with motor impairments.
- Keywords: AAC, gesture recognition, motor impairments, participatory design, machine learning, Transformer models, self-supervised learning, multimodal data, personalized systems, accessibility.
Background and Problem
- Problem / challenge: Existing AAC systems are either aided (relying on external devices) or unaided (using body-based gestures), but these approaches are traditionally treated as mutually exclusive. Aided AAC offers standardized output but is slow and cumbersome for users with motor impairments, while unaided AAC is natural but lacks intelligibility with unfamiliar partners. There is a lack of systems that integrate the benefits of both approaches.
- Significance: Combining aided and unaided AAC could enable faster, more natural communication while maintaining intelligibility, particularly for individuals with motor impairments who face significant barriers in using traditional AAC systems.
- Motivation and related work: Prior work has explored gesture recognition and machine learning for AAC, but these efforts often rely on able-bodied datasets and predefined gestures, failing to account for the idiosyncratic and variable movements of motor-impaired users. This paper aims to address these gaps by developing a personalized, adaptive system.
Solution
- Proposed approach: AllyAAC, a wearable system combining wrist-worn IMU sensors with a smartphone app, enabling real-time recognition of personalized gestures and converting them into synthesized speech.
- Novelty:
- Design principles for integrating aided and unaided AAC based on 18 months of participatory design with AAC users.
- A multimodal dataset of over 600,000 data points featuring atypical gestures paired with IMU data.
- Development of personalized Transformer-based gesture recognition models with self-supervised pretraining strategies.
- A semi-automatic annotation pipeline reducing annotation time by 67%.
- Procedure and key techniques:
- Participatory design sessions with AAC users to identify challenges and opportunities.
- Iterative development of AllyAAC, including gesture recording, annotation, and model training.
- Use of self-supervised learning (contrastive learning, CPC, masked reconstruction) to pretrain models on unlabeled data.
- Deployment of personalized models for real-time gesture recognition and speech output.
Results
- Concrete findings:
- The semi-automatic annotation pipeline achieved an F1-score of 0.95 and reduced annotation time by 67%.
- Large Transformer-based models achieved a mean F1-score of 0.871, a significant improvement over the baseline rule-based model (F1-score of 0.589).
- Human evaluation confirmed a mean precision of 0.88 for large models, with high inter-rater agreement (AC1 = 0.92).
- Advantage over baselines: The large models outperformed the baseline by an average of 28% in F1-score, particularly excelling in recognizing complex or overlapping gestures.
- Experiments / evaluation:
- Field study with 14 participants (8 motor-impaired, 6 non-disabled) at a conference.
- Quantitative evaluation of models using F1-scores and human ratings.
- Analysis of pretraining strategies, showing different strategies excel for different gesture types.
- Limitations and future work:
- Recognition accuracy declined with overlapping gestures or increased vocabulary size.
- Sensor placement and calibration challenges remain.
- Future work includes exploring multimodal sensing (e.g., combining IMU and vision), automating data transfer, and developing context-aware gesture recognition.
Summary
This paper introduces AllyAAC, a system that integrates aided and unaided AAC by recognizing personalized gestures and converting them into synthesized speech. Developed through 18 months of participatory design with AAC users, the system addresses challenges in gesture recognition for motor-impaired individuals using self-supervised learning and a semi-automatic annotation pipeline. Quantitative and qualitative evaluations demonstrate significant improvements in recognition accuracy and usability. The work contributes a novel dataset, adaptive models, and design principles, paving the way for more inclusive and personalized AAC technologies.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)