From User Perceptions to Technical Improvement: Enabling People Who Stutter to Better Use Speech Recognition

Intelligent Voice Assistants (Alexa, Siri, etc.)Voice Accessibility

Title of the Paper

From User Perceptions to Technical Improvement: Enabling People Who Stutter to Better Use Speech Recognition

Paper Information

  • Field of Study: Human-Computer Interaction, Speech Recognition, Assistive Technology, and Accessibility
  • Keywords: Speech Input, Accessibility, Stuttering, Voice Assistants, Speech Transcription

Research Background and Issues

  • Problems and Challenges: Current consumer-grade speech recognition systems perform poorly when processing speech input from people who stutter, including issues such as speech truncation, recognition errors, and transcription results that fail to reflect user intent.
  • Significance: Speech recognition systems are widely used in daily life, but their lack of accessibility makes it difficult for people who stutter to use voice assistants and transcription services, further hindering their participation in the digital society.
  • Motivation and Related Work: The authors observed that current improvements in speech technology for people who stutter are mostly focused on optimizing automatic speech recognition (ASR) models, with a lack of in-depth research based on user experience. Previous studies have preliminarily revealed that time constraints in voice assistants and social pressure may affect the willingness of people who stutter to use these systems, but they have neither quantified the performance deficiencies of speech technologies nor explored related technical solutions.

Solution

  • Proposed Methods:

    1. Endpoint Detection Model Optimization: Adjusting the truncation threshold of the model based on the speech characteristics of people who stutter to reduce premature truncation of speech.
    2. ASR Decoder Parameter Optimization: Tuning decoder settings to reduce recognition errors caused by stuttered speech.
    3. Post-Processing for Speech Flow Optimization: Improving transcription results by addressing stuttering phenomena such as repeated words or filler words through post-processing algorithms.
  • Innovativeness: The proposed solutions can be lightweightly integrated into existing speech recognition systems, avoiding the time-consuming and data-intensive retraining of ASR models. These techniques improve the adaptability of the systems to stuttered speech. Additionally, the study bridges the gap between technical improvements and user experience by combining quantitative performance analysis with subjective user experience surveys.

  • Implementation Steps and Techniques:

    1. Using a large dataset of speech commands and transcriptions recorded by people who stutter, adjust the endpoint detection model's threshold to balance truncation rate and system response delay.
    2. Optimize the speech decoder by increasing the weight of the language model, reducing speech insertion errors, and task-specific tuning.
    3. Apply a language model for post-processing transcription outputs, filtering repeated words and removing potential filler words.

Research Outcomes

  • Specific Results:

    • For endpoint detection: By optimizing the threshold, the speech truncation rate significantly decreased from a baseline of 23.8% to 2.5% (for moderate stutterers), while maintaining an acceptable response delay (1.7 seconds).
    • For ASR performance: Decoder tuning reduced the average word error rate (WER) of transcriptions from 25.4% to 12.4%; combined with speech flow optimization, WER was further reduced to 9.9%.
    • For speech flow optimization: Post-processing of repeated words and filler words significantly improved transcription accuracy, with approximately 64.7% of repeated words successfully corrected and minimal negative impact.
    • Historical trend analysis showed that while speech recognition models have improved in handling stuttered speech over the past five years, further optimization specific to stuttering characteristics is still needed.
  • Comparative Advantages: Compared to existing speech systems, the proposed optimization significantly improved recognition accuracy for people who stutter when using voice assistants and transcription systems (IER reduced from 10.4% to 5.4%), and the algorithms are easy to integrate into existing systems.

  • Experimental or Evaluation Results:

    • Experimental data showed that different speech flow characteristics of people who stutter significantly impact recognition accuracy. For example, word insertion errors are strongly correlated with partial word repetitions.
    • The combination of the three technical improvements maximized performance gains, particularly benefiting moderate to severe stutterers.
  • Limitations and Future Directions:

    • Limitations: Data collection was biased toward quiet personal environments, lacking the complexity of interactions in social contexts. The applicability of speech optimization models may vary with different speech tasks. Additionally, some users may be reluctant to actively choose language correction features.
    • Future Directions: Develop ASR models that can adaptively adjust based on real-time user feedback, validate these technical improvements in voice assistants to assess their impact on real-world interactions, explore broader applicability to other speech disorders (e.g., dysarthria, aphasia), and optimize interaction experiences in social contexts.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/95758/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581224
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
9 authors
sell
Subtopics
Intelligent Voice Assistants (Alexa, Siri, etc.), Voice Accessibility
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
5 related papers