Challenges in Automatic Speech Recognition for Adults with Cognitive Impairment
Authors
Paper Title
Challenges in Automatic Speech Recognition for Adults with Cognitive Impairment
Publication Info
- Topic area: Automatic speech recognition (ASR) performance for older adults with cognitive impairment.
- Keywords: ASR, cognitive impairment, dementia, mild cognitive impairment (MCI), voice assistants, acoustic analysis, human-computer interaction (HCI), aging-in-place, ability-based design, Whisper.
Background and Problem
- Problem / challenge: Automatic speech recognition systems show reduced accuracy for older adults with cognitive impairment, particularly those with dementia, but the underlying acoustic factors contributing to these disparities are poorly understood.
- Significance: ASR errors can hinder the usability of voice-enabled AgeTech systems, impacting trust, satisfaction, and adoption, especially for older adults with cognitive impairment who rely on these technologies for daily living support.
- Motivation and related work: Prior research has identified lower ASR accuracy for older adults and individuals with neurological conditions but has focused primarily on descriptive statistics or ASR errors as inputs for cognitive impairment detection. This study aims to quantify ASR disparities and analyze acoustic features to inform inclusive design.
Solution
- Proposed approach: Evaluation of ASR performance using Whisper models for older adults across cognitive groups, combined with acoustic analysis of speech features to identify predictors of transcription accuracy.
- Novelty:
- Empirical evidence showing reduced ASR accuracy for older adults with dementia compared to cognitively normal controls.
- Acoustic analysis linking speech characteristics (e.g., intensity, shimmer, pause ratio) to ASR error rates.
- Design implications for personalized ASR models, human-in-the-loop correction, and interaction-level adaptations for voice-based AgeTech systems.
- Procedure and key techniques:
- Analysis of 83 older adults reading pre-scripted voice commands to an Amazon Echo device.
- Comparison of ASR performance using Whisper ‘small’ and ‘medium’ models.
- Acoustic feature extraction (e.g., speech rate, articulation rate, shimmer, intensity) using Praat.
- Mixed-effects logistic regression to model word error rate (WER) and identify predictors of ASR accuracy.
Results
- Concrete findings:
- Adults with dementia had significantly higher WER (∼0.21) compared to cognitively normal adults (∼0.08), with no significant difference between cognitively normal and MCI groups.
- Acoustic features such as lower intensity, higher shimmer, and reduced pause ratio were associated with higher WER.
- Whisper ‘medium’ model reduced WER overall, with greater improvements for individuals with dementia.
- Advantage over baselines:
- Whisper ‘medium’ model demonstrated lower WER compared to the ‘small’ model, particularly for individuals with dementia and IoT commands.
- Experiments / evaluation:
- Dataset: Voice Assistant System (VAS) corpus with 83 participants categorized as cognitively normal, MCI, or dementia.
- Metrics: Word error rate (WER) for ASR accuracy; acoustic feature analysis.
- Controlled task: Reading pre-scripted voice commands across categories (IoT, reminders, questions).
- Limitations and future work:
- Focus on scripted speech limits generalization to spontaneous speech interactions.
- Lack of clinical diagnoses for MCI and dementia etiology in the dataset.
- Study restricted to English speakers in the U.S.; findings may not generalize to other languages or cultural contexts.
- Future research needed on spontaneous speech, cross-linguistic comparisons, and adaptation for low-resource languages.
Summary
This study demonstrates that ASR accuracy is significantly lower for older adults with dementia compared to cognitively normal adults, with disparities linked to acoustic features such as intensity, shimmer, and pause ratio. The Whisper ‘medium’ model showed improved transcription accuracy, particularly for individuals with dementia. These findings highlight the need for speaker-personalized ASR models, human-in-the-loop correction systems, and interaction-level adaptations to support equitable access to voice-based AgeTech systems. Future work should explore spontaneous speech, clinical diagnoses, and cross-linguistic applications to further address ASR disparities for individuals with cognitive impairment.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 67%
"I hope I never need one": Unpacking Stigma in Aging in Place Technology
CHI '22· Aging-Friendly Technology Design +1
- 67%
Invisibility or Visibility in Intimate Care at the Workplace? Examining the Use of Breast Pumps
CHI '23· Aging-Friendly Technology Design +1
- 67%
Hidden Opportunities for Elder Living: Understanding Shared Technology Troubles and Benefits for Older Adults in the UK Cost of Living Crisis
CHI '25· Aging-Friendly Technology Design +1
- 67%
Values Across Contexts: Understanding How Older Adults Enact What Matters Through Technology
CHI '26· Aging-Friendly Technology Design +1
Based on Jaccard similarity of research subtopics & professions (≥60%)