Unspoken Sound: Identifying Trends in Non-Speech Audio Captioning on YouTube

Multilingual & Cross-Cultural Voice InteractionVoice AccessibilityDeaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)Speech-Language Pathologists & AudiologistsContent Creators (YouTubers, Podcasters)

Title of the Paper

Unspoken Sound: Identifying Trends in Non-Speech Audio Captioning on YouTube

Paper Information

  • Subject Area: Research on Non-Speech Audio Information (NSI) in Closed Captions for Video Content
  • Keywords: Dataset, Closed Captions, Non-Speech Information (NSI), Extra Speech Information (ESI), Caption Quality, Audio Description, Video Accessibility

Research Background and Problem Statement

  • Issues or Challenges Identified by the Authors:

    • Insufficient caption coverage of non-speech audio information (NSI) in online videos, particularly on YouTube.
    • Automatic Speech Recognition (ASR) technology supports only limited categories of NSI (e.g., "[Laughter]", "[Music]", "[Applause]") and lacks support for other types of NSI.
    • A low proportion of manually generated captions, with a large number of videos lacking any form of captions, especially in content produced by major studios.
    • The lack of data and analytical tools makes it difficult to study the quality of NSI captions online.
  • Why This Problem is Important:

    • Video content accessibility is critical for d/Deaf and Hard of Hearing (DHH) audiences, yet current captioning systems exhibit significant deficiencies in the quality and implementation of NSI information.
    • NSI is essential for understanding the narrative and emotional aspects of video content. Without captions, these audiences may miss important information.
    • The vast volume of video uploads on platforms like YouTube highlights the need to improve NSI coverage to enhance global video content accessibility.
  • Motivation and Related Work:

    • Limited analysis exists on the current state of NSI, with most research focusing on linguistic information or video user experience.
    • Related work has explored NSI categorization, automated audio description (AAC) models, and video caption user preferences, but lacks in-depth investigation of NSI implementation on popular video platforms.

Solution

  • Research Methods:

    • Created a dataset covering 700,000+ YouTube videos, analyzing their caption content and metadata (spanning 2013 to 2022, including popular and major studio samples).
    • Manually annotated 1,799 videos from three target years (2013, 2018, 2022), identifying various types of NSI.
    • Proposed an "NSI Estimator" to automatically detect NSI information in captions, based on common NSI symbols and formats (e.g., "[", "]", ":").
    • Defined metrics to evaluate NSI caption quality (e.g., NSI density CPMIP and NSI presence ratio).
  • Innovations:

    • Combined manual annotation and automated estimation to comprehensively assess the state of NSI in YouTube captions, revealing long-term trends in platform captioning practices through dataset creation.
    • Introduced new quality evaluation metrics for NSI and validated a machine-assisted NSI estimation method.

Research Findings

  • Key Findings:

    • Automatic captions dominate, but there is still a significant lack of NSI descriptions. The proportion of manual captions is only 6-8%.
    • Shorter videos have higher NSI caption density, while NSI captions are significantly reduced in longer videos.
    • There are notable differences in NSI category implementation: major studio samples focus primarily on extra speech information (e.g., speaker identification), while popular samples cover a broader range of NSI types.
    • Video topic and duration significantly affect the frequency and quality of NSI, with certain categories (e.g., music, religious themes) having richer NSI captions.
  • Comparison with Existing Solutions:

    • Provides a more comprehensive data analysis framework than existing audio description systems, identifying specific areas where captioning technology fails to meet the needs of DHH audiences.
    • Offers a more detailed examination of how video topics and production processes influence NSI captions compared to prior studies.
  • Experimental or Evaluation Results:

    • The dataset shows an NSI presence ratio of 88-92%, but the density is significantly lower, far below the 3 CPMIP recommended by high-quality captioning references such as Zdenek et al.
    • Diversity entropy analysis of major studio samples indicates lower NSI coverage and category diversity compared to popular samples.
  • Limitations and Future Directions:

    • The analysis is limited to English-based captions and content primarily from the United States, which may not reflect the global state of YouTube videos.
    • The lack of further detailed exploration into caption generation and media processing workflows leaves some trends' specific mechanisms unexplained.
    • Future work should develop quality evaluation standards for NSI, enhance the descriptive capabilities of automated models, and explore the potential of community-driven caption editing tools to improve NSI.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/146864/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642162
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
8 authors
sell
Subtopics
Multilingual & Cross-Cultural Voice Interaction, Voice Accessibility, Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)
work
Professions
Speech-Language Pathologists & Audiologists, Content Creators (YouTubers, Podcasters)
article
Content Status
Full text indexed
hub
Related Papers
1 related papers