Like, Comment & Caption: A Decade of Social Media Video Caption Research (2015–2025)

Honorable Mention
Voice AccessibilitySocial Platform Design & User BehaviorDeaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)Universal & Inclusive DesignPedestrians & Vulnerable Road UsersUI/UX DesignersContent Creators (YouTubers, Podcasters)Speech-Language Pathologists & Audiologists

Paper Title

Like, Comment & Caption: A Decade of Social Media Video Caption Research (2015–2025)

Publication Info

  • Topic area: Systematic review of social media video captioning systems and practices.
  • Keywords: social media video captions, accessibility, participatory captioning, deaf and hard of hearing (DHH), neurodivergent viewers, creator economy, platform design, non-speech information, automatic captions, user-generated captions.

Background and Problem

  • Problem / challenge: Social media video captioning systems lack a systematic, design-oriented synthesis, leading to interventions that overlook platform constraints and accessibility needs.
  • Significance: Captions are critical for accessibility, engagement, and visibility on social media platforms, especially for DHH, neurodivergent, and multilingual users.
  • Motivation and related work: Prior research has explored captions in television, streaming, and videoconferencing, but social media platforms introduce unique challenges due to their participatory, algorithmic, and creator-driven nature. Existing studies are fragmented, focusing on specific platforms, communities, or technical aspects without a comprehensive synthesis.

Solution

  • Proposed approach: The paper introduces the framework of Participatory Captioning, emphasizing the co-production of captions by viewers, creators, and platforms.
  • Novelty:
    1. Systematic review of 36 papers (2015–2025) across HCI, accessibility, media studies, education, and language learning.
    2. Identification of four key themes in social media video captioning (SMVC): caption types, viewer perspectives, creator perspectives, and technical systems/datasets.
    3. Proposal of Participatory Captioning as a collaborative framework for SMVC design and research.
    4. Design implications and future research directions for improving SMVC systems.
  • Procedure and key techniques:
    • Conducted a systematic review using a two-phase search strategy on Google Scholar and top HCI venues.
    • Coded papers on dimensions such as stakeholder roles, community types, data collection methods, and contributions.
    • Synthesized findings into themes and proposed design opportunities for participatory captioning systems.

Results

  • Concrete findings:
    • SMVC research has grown significantly since 2019, with increasing focus on multi-platform and multi-community studies.
    • Four SMVC types identified: automatic captions, user-generated captions, non-speech information, and captions for sign language.
    • Persistent challenges include caption accuracy, synchronization, non-speech information coverage, and inadequate support for sign language.
    • Viewers value captions for accessibility, learning, and cultural connection but face barriers such as missing captions and limited customization.
    • Creators are motivated by accessibility, branding, and engagement but struggle with labor-intensive captioning processes and inadequate tools.
  • Advantage over baselines:
    • Highlights gaps in current SMVC systems, such as lack of participatory feedback mechanisms and culturally grounded captioning practices.
    • Proposes Participatory Captioning as a novel framework to address these gaps by involving viewers and creators in caption co-production.
  • Experiments / evaluation:
    • Reviewed 36 peer-reviewed papers using qualitative coding and thematic analysis.
    • Identified trends in platform and community diversity, methodological approaches, and key challenges in SMVC research.
  • Limitations and future work:
    • Limited to English-language studies, potentially missing insights from non-English contexts.
    • Focused on Google Scholar and HCI venues, which may exclude broader interdisciplinary work.
    • Future research should expand to underrepresented communities, non-Western platforms, and emerging technologies like multimodal AI.

Summary

This paper systematically reviews a decade of research on social media video captioning systems, identifying key themes, challenges, and opportunities. It introduces Participatory Captioning, a framework emphasizing the co-production of captions by viewers, creators, and platforms to address accessibility and engagement gaps. The findings highlight the need for improved tools, feedback systems, and culturally inclusive practices. Future research should expand to diverse communities, platforms, and technologies, ensuring that SMVC systems evolve to meet the needs of all users.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223393/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791868
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
Honorable Mention
group
Authors
7 authors
sell
Subtopics
Voice Accessibility, Social Platform Design & User Behavior, Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration), Universal & Inclusive Design
work
Professions
Pedestrians & Vulnerable Road Users, UI/UX Designers, Content Creators (YouTubers, Podcasters), Speech-Language Pathologists & Audiologists
article
Content Status
Full text indexed
hub
Related Papers
1 related papers