Lost in Transcription: Subtitle Errors in Automatic Speech Recognition Reduce Speaker and Content Evaluations
Authors
Paper Title
Lost in Transcription: Subtitle Errors in Automatic Speech Recognition Reduce Speaker and Content Evaluations
Publication Info
- Topic area: Impact of ASR subtitle errors on speaker and content evaluations.
- Keywords: ASR, subtitle quality, speaker evaluation, content evaluation, accent bias, subtitle errors, audience perception, automatic transcription, cognitive load, inclusivity.
Background and Problem
- Problem / challenge: ASR systems often produce error-prone subtitles, with biases in performance across demographic groups. These errors may negatively affect how audiences evaluate speakers and their content, particularly for speakers with non-standard accents.
- Significance: Subtitle errors can exacerbate existing social biases, potentially disadvantaging speakers in professional, educational, and social contexts.
- Motivation and related work: Prior research has shown that ASR systems exhibit performance disparities based on race, gender, and accent, leading to emotional and professional harms. However, existing studies have not fully controlled for confounding factors like speaker identity and appearance, leaving gaps in understanding the direct effects of subtitle errors on speaker and content evaluations.
Solution
- Proposed approach: A mixed-factorial online experiment to assess the impact of subtitle errors and speaker accents on audience evaluations of speakers and their content.
- Novelty:
- Controlled for speaker identity, appearance, and delivery using AI-generated accents.
- Examined the combined effects of subtitle quality and speaker accent on evaluations.
- Developed realistic error-prone subtitles using real-world ASR systems.
- Introduced a novel methodological design to isolate the effects of subtitle errors.
- Procedure and key techniques:
- Participants (N=207) watched two videos featuring speakers with either Standard American English (SAE) or non-SAE accents.
- Each participant viewed one video with accurate subtitles and one with error-prone subtitles.
- Evaluations were collected on subtitle quality, speaker delivery, and content quality.
- AI-generated audio tracks ensured consistent speaker characteristics across accent conditions.
Results
- Concrete findings:
- Subtitle errors significantly reduced speaker evaluations (β = 0.451, p < 0.001) and content evaluations (β = 0.341, p < 0.001).
- Accurate subtitles led to higher evaluations for both SAE and non-SAE speakers (e.g., Mnon-SAE=3.98 vs. Mnon-SAE=3.48 for speaker evaluations).
- No significant differences in evaluations based on speaker accent (β = −0.056, p = 0.604 for speaker evaluations; β = −0.069, p = 0.512 for content evaluations).
- Advantage over baselines: Controlled for confounding factors like speaker appearance and delivery, enabling clearer attribution of effects to subtitle quality.
- Experiments / evaluation:
- Videos were sourced from TED Talks and manipulated using AI-generated audio to create SAE and non-SAE accent conditions.
- Error-prone subtitles were generated using Google Meet, which had the highest word error rate (WER = 0.31).
- Measures included subtitle quality, speaker delivery, and content quality, analyzed using linear mixed models.
- Limitations and future work:
- Limited to South Asian male speakers; findings may not generalize across demographics.
- Participants were primarily from WEIRD populations, which may not reflect broader audience perceptions.
- Did not investigate mechanisms behind lower evaluations or whether errors were attributed to the speaker or the system.
- Future work should explore perceived subtitle quality, error significance, and adaptive subtitle designs.
Summary
This study demonstrates that ASR-generated subtitle errors significantly harm speaker and content evaluations, regardless of speaker accent. By controlling for speaker identity and using AI-generated accents, the research isolates the effects of subtitle quality, showing that accurate subtitles improve audience perceptions. While no accent-based disparities were detected, the findings highlight the potential for ASR errors to compound existing biases in professional and social contexts. Future research should explore adaptive subtitle designs and audience perceptions of subtitle quality to mitigate these harms and foster more inclusive ASR systems.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)