Challenges in Automatic Speech Recognition for Adults with Cognitive Impairment

Voice AccessibilityAging-Friendly Technology DesignAging-in-Place Assistance SystemsPhysicians, Nurses & CliniciansElderly Care WorkersFamily Caregivers

Paper Title

Challenges in Automatic Speech Recognition for Adults with Cognitive Impairment

Publication Info

  • Topic area: Automatic speech recognition (ASR) performance for older adults with cognitive impairment.
  • Keywords: ASR, cognitive impairment, dementia, mild cognitive impairment (MCI), voice assistants, acoustic analysis, human-computer interaction (HCI), aging-in-place, ability-based design, Whisper.

Background and Problem

  • Problem / challenge: Automatic speech recognition systems show reduced accuracy for older adults with cognitive impairment, particularly those with dementia, but the underlying acoustic factors contributing to these disparities are poorly understood.
  • Significance: ASR errors can hinder the usability of voice-enabled AgeTech systems, impacting trust, satisfaction, and adoption, especially for older adults with cognitive impairment who rely on these technologies for daily living support.
  • Motivation and related work: Prior research has identified lower ASR accuracy for older adults and individuals with neurological conditions but has focused primarily on descriptive statistics or ASR errors as inputs for cognitive impairment detection. This study aims to quantify ASR disparities and analyze acoustic features to inform inclusive design.

Solution

  • Proposed approach: Evaluation of ASR performance using Whisper models for older adults across cognitive groups, combined with acoustic analysis of speech features to identify predictors of transcription accuracy.
  • Novelty:
    1. Empirical evidence showing reduced ASR accuracy for older adults with dementia compared to cognitively normal controls.
    2. Acoustic analysis linking speech characteristics (e.g., intensity, shimmer, pause ratio) to ASR error rates.
    3. Design implications for personalized ASR models, human-in-the-loop correction, and interaction-level adaptations for voice-based AgeTech systems.
  • Procedure and key techniques:
    • Analysis of 83 older adults reading pre-scripted voice commands to an Amazon Echo device.
    • Comparison of ASR performance using Whisper ‘small’ and ‘medium’ models.
    • Acoustic feature extraction (e.g., speech rate, articulation rate, shimmer, intensity) using Praat.
    • Mixed-effects logistic regression to model word error rate (WER) and identify predictors of ASR accuracy.

Results

  • Concrete findings:
    • Adults with dementia had significantly higher WER (∼0.21) compared to cognitively normal adults (∼0.08), with no significant difference between cognitively normal and MCI groups.
    • Acoustic features such as lower intensity, higher shimmer, and reduced pause ratio were associated with higher WER.
    • Whisper ‘medium’ model reduced WER overall, with greater improvements for individuals with dementia.
  • Advantage over baselines:
    • Whisper ‘medium’ model demonstrated lower WER compared to the ‘small’ model, particularly for individuals with dementia and IoT commands.
  • Experiments / evaluation:
    • Dataset: Voice Assistant System (VAS) corpus with 83 participants categorized as cognitively normal, MCI, or dementia.
    • Metrics: Word error rate (WER) for ASR accuracy; acoustic feature analysis.
    • Controlled task: Reading pre-scripted voice commands across categories (IoT, reminders, questions).
  • Limitations and future work:
    • Focus on scripted speech limits generalization to spontaneous speech interactions.
    • Lack of clinical diagnoses for MCI and dementia etiology in the dataset.
    • Study restricted to English speakers in the U.S.; findings may not generalize to other languages or cultural contexts.
    • Future research needed on spontaneous speech, cross-linguistic comparisons, and adaptation for low-resource languages.

Summary

This study demonstrates that ASR accuracy is significantly lower for older adults with dementia compared to cognitively normal adults, with disparities linked to acoustic features such as intensity, shimmer, and pause ratio. The Whisper ‘medium’ model showed improved transcription accuracy, particularly for individuals with dementia. These findings highlight the need for speaker-personalized ASR models, human-in-the-loop correction systems, and interaction-level adaptations to support equitable access to voice-based AgeTech systems. Future work should explore spontaneous speech, clinical diagnoses, and cross-linguistic applications to further address ASR disparities for individuals with cognitive impairment.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222053/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791581
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Voice Accessibility, Aging-Friendly Technology Design, Aging-in-Place Assistance Systems
work
Professions
Physicians, Nurses & Clinicians, Elderly Care Workers, Family Caregivers
article
Content Status
Full text indexed
hub
Related Papers
4 related papers