From User Perceptions to Technical Improvement: Enabling People Who Stutter to Better Use Speech Recognition
Authors
Title of the Paper
From User Perceptions to Technical Improvement: Enabling People Who Stutter to Better Use Speech Recognition
Paper Information
- Field of Study: Human-Computer Interaction, Speech Recognition, Assistive Technology, and Accessibility
- Keywords: Speech Input, Accessibility, Stuttering, Voice Assistants, Speech Transcription
Research Background and Issues
- Problems and Challenges: Current consumer-grade speech recognition systems perform poorly when processing speech input from people who stutter, including issues such as speech truncation, recognition errors, and transcription results that fail to reflect user intent.
- Significance: Speech recognition systems are widely used in daily life, but their lack of accessibility makes it difficult for people who stutter to use voice assistants and transcription services, further hindering their participation in the digital society.
- Motivation and Related Work: The authors observed that current improvements in speech technology for people who stutter are mostly focused on optimizing automatic speech recognition (ASR) models, with a lack of in-depth research based on user experience. Previous studies have preliminarily revealed that time constraints in voice assistants and social pressure may affect the willingness of people who stutter to use these systems, but they have neither quantified the performance deficiencies of speech technologies nor explored related technical solutions.
Solution
-
Proposed Methods:
- Endpoint Detection Model Optimization: Adjusting the truncation threshold of the model based on the speech characteristics of people who stutter to reduce premature truncation of speech.
- ASR Decoder Parameter Optimization: Tuning decoder settings to reduce recognition errors caused by stuttered speech.
- Post-Processing for Speech Flow Optimization: Improving transcription results by addressing stuttering phenomena such as repeated words or filler words through post-processing algorithms.
-
Innovativeness: The proposed solutions can be lightweightly integrated into existing speech recognition systems, avoiding the time-consuming and data-intensive retraining of ASR models. These techniques improve the adaptability of the systems to stuttered speech. Additionally, the study bridges the gap between technical improvements and user experience by combining quantitative performance analysis with subjective user experience surveys.
-
Implementation Steps and Techniques:
- Using a large dataset of speech commands and transcriptions recorded by people who stutter, adjust the endpoint detection model's threshold to balance truncation rate and system response delay.
- Optimize the speech decoder by increasing the weight of the language model, reducing speech insertion errors, and task-specific tuning.
- Apply a language model for post-processing transcription outputs, filtering repeated words and removing potential filler words.
Research Outcomes
-
Specific Results:
- For endpoint detection: By optimizing the threshold, the speech truncation rate significantly decreased from a baseline of 23.8% to 2.5% (for moderate stutterers), while maintaining an acceptable response delay (1.7 seconds).
- For ASR performance: Decoder tuning reduced the average word error rate (WER) of transcriptions from 25.4% to 12.4%; combined with speech flow optimization, WER was further reduced to 9.9%.
- For speech flow optimization: Post-processing of repeated words and filler words significantly improved transcription accuracy, with approximately 64.7% of repeated words successfully corrected and minimal negative impact.
- Historical trend analysis showed that while speech recognition models have improved in handling stuttered speech over the past five years, further optimization specific to stuttering characteristics is still needed.
-
Comparative Advantages: Compared to existing speech systems, the proposed optimization significantly improved recognition accuracy for people who stutter when using voice assistants and transcription systems (IER reduced from 10.4% to 5.4%), and the algorithms are easy to integrate into existing systems.
-
Experimental or Evaluation Results:
- Experimental data showed that different speech flow characteristics of people who stutter significantly impact recognition accuracy. For example, word insertion errors are strongly correlated with partial word repetitions.
- The combination of the three technical improvements maximized performance gains, particularly benefiting moderate to severe stutterers.
-
Limitations and Future Directions:
- Limitations: Data collection was biased toward quiet personal environments, lacking the complexity of interactions in social contexts. The applicability of speech optimization models may vary with different speech tasks. Additionally, some users may be reluctant to actively choose language correction features.
- Future Directions: Develop ASR models that can adaptively adjust based on real-time user feedback, validate these technical improvements in voice assistants to assess their impact on real-world interactions, explore broader applicability to other speech disorders (e.g., dysarthria, aphasia), and optimize interaction experiences in social contexts.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- Why do current speech recognition systems struggle to accurately process speech input from people who stutter?Category: XR Cybersickness Detection and MitigationSimilar questionsarrow_forward
- How can existing speech recognition systems be optimized to better support people who stutter?Category: Voice Assistant General Design and User ExperienceSimilar questionsarrow_forward
- How do improvements in speech optimization technologies affect the user experience of people who stutter?Category: Voice Assistant General Design and User ExperienceSimilar questionsarrow_forward
Practical Problems
1- People who stutter face barriers to digital social participation when using voice assistants and transcription services.Category: Voice Assistant General Design and User ExperienceSimilar questionsarrow_forward
- 67%
A Large Inclusive Study of Human Listening Rates
CHI '18· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 67%
“Accessibility Came by Accident”: Use of Voice-Controlled Intelligent Personal Assistants by People with Disabilities
CHI '18· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 67%
WESPER: Zero-shot and Realtime Whisper to Normal Voice Conversion for Whisper-based Speech interactions
CHI '23· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 67%
A Mixed-Methods Approach to Understanding User Trust after Voice Assistant Failures
CHI '23· Intelligent Voice Assistants (Alexa, Siri, etc.) +2
- 67%
Morae: Proactively Pausing UI Agents for User Choices
UIST '25· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)