Unspoken Sound: Identifying Trends in Non-Speech Audio Captioning on YouTube
Authors
Multilingual & Cross-Cultural Voice InteractionVoice AccessibilityDeaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)Speech-Language Pathologists & AudiologistsContent Creators (YouTubers, Podcasters)
Title of the Paper
Unspoken Sound: Identifying Trends in Non-Speech Audio Captioning on YouTube
Paper Information
- Subject Area: Research on Non-Speech Audio Information (NSI) in Closed Captions for Video Content
- Keywords: Dataset, Closed Captions, Non-Speech Information (NSI), Extra Speech Information (ESI), Caption Quality, Audio Description, Video Accessibility
Research Background and Problem Statement
-
Issues or Challenges Identified by the Authors:
- Insufficient caption coverage of non-speech audio information (NSI) in online videos, particularly on YouTube.
- Automatic Speech Recognition (ASR) technology supports only limited categories of NSI (e.g., "[Laughter]", "[Music]", "[Applause]") and lacks support for other types of NSI.
- A low proportion of manually generated captions, with a large number of videos lacking any form of captions, especially in content produced by major studios.
- The lack of data and analytical tools makes it difficult to study the quality of NSI captions online.
-
Why This Problem is Important:
- Video content accessibility is critical for d/Deaf and Hard of Hearing (DHH) audiences, yet current captioning systems exhibit significant deficiencies in the quality and implementation of NSI information.
- NSI is essential for understanding the narrative and emotional aspects of video content. Without captions, these audiences may miss important information.
- The vast volume of video uploads on platforms like YouTube highlights the need to improve NSI coverage to enhance global video content accessibility.
-
Motivation and Related Work:
- Limited analysis exists on the current state of NSI, with most research focusing on linguistic information or video user experience.
- Related work has explored NSI categorization, automated audio description (AAC) models, and video caption user preferences, but lacks in-depth investigation of NSI implementation on popular video platforms.
Solution
-
Research Methods:
- Created a dataset covering 700,000+ YouTube videos, analyzing their caption content and metadata (spanning 2013 to 2022, including popular and major studio samples).
- Manually annotated 1,799 videos from three target years (2013, 2018, 2022), identifying various types of NSI.
- Proposed an "NSI Estimator" to automatically detect NSI information in captions, based on common NSI symbols and formats (e.g., "[", "]", ":").
- Defined metrics to evaluate NSI caption quality (e.g., NSI density CPMIP and NSI presence ratio).
-
Innovations:
- Combined manual annotation and automated estimation to comprehensively assess the state of NSI in YouTube captions, revealing long-term trends in platform captioning practices through dataset creation.
- Introduced new quality evaluation metrics for NSI and validated a machine-assisted NSI estimation method.
Research Findings
-
Key Findings:
- Automatic captions dominate, but there is still a significant lack of NSI descriptions. The proportion of manual captions is only 6-8%.
- Shorter videos have higher NSI caption density, while NSI captions are significantly reduced in longer videos.
- There are notable differences in NSI category implementation: major studio samples focus primarily on extra speech information (e.g., speaker identification), while popular samples cover a broader range of NSI types.
- Video topic and duration significantly affect the frequency and quality of NSI, with certain categories (e.g., music, religious themes) having richer NSI captions.
-
Comparison with Existing Solutions:
- Provides a more comprehensive data analysis framework than existing audio description systems, identifying specific areas where captioning technology fails to meet the needs of DHH audiences.
- Offers a more detailed examination of how video topics and production processes influence NSI captions compared to prior studies.
-
Experimental or Evaluation Results:
- The dataset shows an NSI presence ratio of 88-92%, but the density is significantly lower, far below the 3 CPMIP recommended by high-quality captioning references such as Zdenek et al.
- Diversity entropy analysis of major studio samples indicates lower NSI coverage and category diversity compared to popular samples.
-
Limitations and Future Directions:
- The analysis is limited to English-based captions and content primarily from the United States, which may not reflect the global state of YouTube videos.
- The lack of further detailed exploration into caption generation and media processing workflows leaves some trends' specific mechanisms unexplained.
- Future work should develop quality evaluation standards for NSI, enhance the descriptive capabilities of automated models, and explore the potential of community-driven caption editing tools to improve NSI.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- What implementation trends exist for non-speech audio information (NSI) in YouTube captions?Category: Captions, Speech Transcription, and Deaf Communication AssistanceSimilar questionsarrow_forward
- How do different video types and producers affect NSI caption coverage and quality?Category: Captions, Speech Transcription, and Deaf Communication AssistanceSimilar questionsarrow_forward
- How can datasets and automated tools improve quality evaluation standards for NSI captions?Category: Captions, Speech Transcription, and Deaf Communication AssistanceSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Deaf users struggle to obtain complete video content from YouTube captions lacking non-speech information.Category: Captions, Speech Transcription, and Deaf Communication AssistanceSimilar questionsarrow_forward
- 60%
Towards AI-driven Sign Language Generation with Non-manual Markers
CHI '25· Voice Accessibility +1
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642162
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
8 authors
sell
Subtopics
Multilingual & Cross-Cultural Voice Interaction, Voice Accessibility, Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)
work
Professions
Speech-Language Pathologists & Audiologists, Content Creators (YouTubers, Podcasters)
article
Content Status
Full text indexed
hub
Related Papers
1 related papers