Lost in Transcription: Subtitle Errors in Automatic Speech Recognition Reduce Speaker and Content Evaluations

Intelligent Voice Assistants (Alexa, Siri, etc.)Voice AccessibilityPrivacy by Design & User ControlAmazon Mechanical Turk WorkersFreelancers (Design, Writing, Translation)HCI Researchers

Paper Title

Lost in Transcription: Subtitle Errors in Automatic Speech Recognition Reduce Speaker and Content Evaluations

Publication Info

  • Topic area: Impact of ASR subtitle errors on speaker and content evaluations.
  • Keywords: ASR, subtitle quality, speaker evaluation, content evaluation, accent bias, subtitle errors, audience perception, automatic transcription, cognitive load, inclusivity.

Background and Problem

  • Problem / challenge: ASR systems often produce error-prone subtitles, with biases in performance across demographic groups. These errors may negatively affect how audiences evaluate speakers and their content, particularly for speakers with non-standard accents.
  • Significance: Subtitle errors can exacerbate existing social biases, potentially disadvantaging speakers in professional, educational, and social contexts.
  • Motivation and related work: Prior research has shown that ASR systems exhibit performance disparities based on race, gender, and accent, leading to emotional and professional harms. However, existing studies have not fully controlled for confounding factors like speaker identity and appearance, leaving gaps in understanding the direct effects of subtitle errors on speaker and content evaluations.

Solution

  • Proposed approach: A mixed-factorial online experiment to assess the impact of subtitle errors and speaker accents on audience evaluations of speakers and their content.
  • Novelty:
    1. Controlled for speaker identity, appearance, and delivery using AI-generated accents.
    2. Examined the combined effects of subtitle quality and speaker accent on evaluations.
    3. Developed realistic error-prone subtitles using real-world ASR systems.
    4. Introduced a novel methodological design to isolate the effects of subtitle errors.
  • Procedure and key techniques:
    • Participants (N=207) watched two videos featuring speakers with either Standard American English (SAE) or non-SAE accents.
    • Each participant viewed one video with accurate subtitles and one with error-prone subtitles.
    • Evaluations were collected on subtitle quality, speaker delivery, and content quality.
    • AI-generated audio tracks ensured consistent speaker characteristics across accent conditions.

Results

  • Concrete findings:
    • Subtitle errors significantly reduced speaker evaluations (β = 0.451, p < 0.001) and content evaluations (β = 0.341, p < 0.001).
    • Accurate subtitles led to higher evaluations for both SAE and non-SAE speakers (e.g., Mnon-SAE=3.98 vs. Mnon-SAE=3.48 for speaker evaluations).
    • No significant differences in evaluations based on speaker accent (β = −0.056, p = 0.604 for speaker evaluations; β = −0.069, p = 0.512 for content evaluations).
  • Advantage over baselines: Controlled for confounding factors like speaker appearance and delivery, enabling clearer attribution of effects to subtitle quality.
  • Experiments / evaluation:
    • Videos were sourced from TED Talks and manipulated using AI-generated audio to create SAE and non-SAE accent conditions.
    • Error-prone subtitles were generated using Google Meet, which had the highest word error rate (WER = 0.31).
    • Measures included subtitle quality, speaker delivery, and content quality, analyzed using linear mixed models.
  • Limitations and future work:
    • Limited to South Asian male speakers; findings may not generalize across demographics.
    • Participants were primarily from WEIRD populations, which may not reflect broader audience perceptions.
    • Did not investigate mechanisms behind lower evaluations or whether errors were attributed to the speaker or the system.
    • Future work should explore perceived subtitle quality, error significance, and adaptive subtitle designs.

Summary

This study demonstrates that ASR-generated subtitle errors significantly harm speaker and content evaluations, regardless of speaker accent. By controlling for speaker identity and using AI-generated accents, the research isolates the effects of subtitle quality, showing that accurate subtitles improve audience perceptions. While no accent-based disparities were detected, the findings highlight the potential for ASR errors to compound existing biases in professional and social contexts. Future research should explore adaptive subtitle designs and audience perceptions of subtitle quality to mitigate these harms and foster more inclusive ASR systems.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222782/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790911
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Intelligent Voice Assistants (Alexa, Siri, etc.), Voice Accessibility, Privacy by Design & User Control
work
Professions
Amazon Mechanical Turk Workers, Freelancers (Design, Writing, Translation), HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers