Title of the Paper

Making Short-Form Videos Accessible with Hierarchical Video Summaries

Paper Information

  • Topic Area: Assistive Technology and Human-Computer Interaction, focusing on the accessibility of short-form videos
  • Keywords: Short-form videos, accessibility, video description, hierarchical summaries, social media, blind users, GPT-4, video visualization

Research Background and Problem

  • Problem or Challenge:

    • Short-form videos (e.g., TikTok, Instagram Reels, and YouTube Shorts) have become major sources of information and entertainment, but they are largely inaccessible to blind and low-vision (BLV) audiences due to rapid visual transitions, dense on-screen text, and background music or meme audio overlays.
    • Existing video description methods (such as embedded audio descriptions or manually added descriptive text) face limitations in short-form videos, struggling to provide timely and synchronized information for fast-paced content.
    • BLV users need reliable ways to assess video accessibility or decide whether to watch, which poses even greater challenges on short-form video platforms.
  • Significance:

    • Improving the accessibility of short-form videos for BLV users enhances their access to mainstream media and promotes social inclusivity.
    • High-quality description systems can address existing technological gaps and provide insights for human-computer interaction research.
  • Research Motivation and Related Work:

    • Previous studies have explored accessibility audio descriptions for long-form videos, but there is a significant gap in research on short-form video accessibility.
    • Hierarchical video summary methods have not been systematically validated in the context of short-form video accessibility.
    • Social media platforms provide limited support for BLV users, who often rely on friends or online communities to supplement descriptive information, which is time-consuming and inefficient.

Solution

  • Method or Solution:

    1. System Design: Developed a system called ShortScribe to provide hierarchical video summaries for BLV users.
    2. Hierarchical Summaries:
      • Brief Description: A summary of fewer than 10 words to help users quickly grasp the video content.
      • Detailed Description: A longer summary offering more semantic details.
      • Shot-by-Shot Description: A step-by-step detailed summary broken down by video segments.
      • Screen Text Extraction: Optical character recognition (OCR) to extract text displayed in the video.
    3. Technical Strategies:
      • Used Google Cloud's ASR for automatic transcription of video audio.
      • Applied scene detection (FFMPEG) to segment the video.
      • Leveraged the BLIP-2 model to generate visual descriptions, combined with GPT-4 for abstraction and integration of data.
  • Innovations:

    • Proposed a hierarchical video summary structure, enabling users to access descriptions with varying levels of detail as needed.
    • Utilized multimodal methods to extract and integrate visual, audio, and operational information from videos.
    • Enhanced the quality of generated video descriptions by incorporating GPT-4's summarization capabilities.
  • Implementation Steps:

    1. Data Processing Pipeline:
      • Videos were segmented into multiple shots, with keyframes computed for each shot.
      • Text extraction (OCR), visual description generation (BLIP-2), and audio transcription were conducted.
    2. Description Generation and Optimization:
      • GPT-4 generated flexible summaries, including brief descriptions, detailed descriptions, and shot-by-shot summaries.
      • Multimodal information (visual, audio, on-screen text) was aggregated to ensure semantic coverage and information accuracy.
    3. Interface Design:
      • Two-tier interface structure: a video player level (displaying brief descriptions) and a description access level.
      • Designed screen reader-friendly features for improved accessibility.

Research Outcomes

  • Specific Results:

    1. ShortScribe significantly improved BLV users' understanding of short-form videos:
      • Users' comprehension scores increased from 2.53 with baseline systems to 5.89 with ShortScribe (out of 7).
      • Users' accuracy in summarizing video content rose from 20% (baseline) to 73%.
    2. User testing revealed specific preferences for different levels of description (e.g., "brief descriptions" for quick browsing, "detailed descriptions" for deeper understanding).
    3. Description coverage rates ranged from 75% (brief descriptions) to 100% (detailed and shot-by-shot descriptions).
  • Comparative Advantages:

    • Compared to existing description technologies, ShortScribe allows users to flexibly choose the appropriate level of information, saving time and improving efficiency in selecting videos of interest.
    • By integrating multimodal data with large language models, the system achieves superior descriptive performance and higher supportability.
  • Experimental or Evaluation Results:

    • In task-based user studies, participants (N=10) explicitly stated that ShortScribe improved their short-form video viewing experience.
    • Technical evaluations showed that most descriptions were accurate, with error rates of 33% for brief descriptions and 57% for detailed descriptions. However, some model misinterpretations (e.g., emotional misjudgments) were observed.
  • Limitations and Future Directions:

    • Limitations:
      1. Certain complex visual actions (e.g., dances) or interactive videos (e.g., reaction videos) were not effectively supported.
      2. Errors in model-generated descriptions affected the accurate understanding of some videos.
      3. Redundant information in the description system (e.g., repetitive content) may impact user experience.
    • Future Directions:
      • Improve the accuracy of vision-to-language models (e.g., adopting improved GPT-V or Google Bard).
      • Explore personalized description generation mechanisms (e.g., user-defined priorities in summaries).
      • Optimize support for complex data generation (e.g., 3D motion reconstruction, multi-layer video descriptions).
      • Extend applications to long-form videos, live-streaming videos, and panoramic video scenarios.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/148297/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642839
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
2 related papers