CrossA11y: Identifying Video Accessibility Issues via Cross-modal Grounding

Best Paper
Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)Universal & Inclusive DesignContent Creators (YouTubers, Podcasters)UI/UX DesignersAssistive Technology Specialists

Document Title

CrossA11y: Identifying Video Accessibility Issues via Cross-modal Grounding

Document Information

  • Subject Area: Video accessibility issue identification, cross-modal technology
  • Keywords: Audio description, closed captions, video, accessibility, cross-modal

Research Background and Issues

  • Identified Issues or Challenges:
    • Videos convey content through visual and auditory information, which poses accessibility challenges for users who cannot access these modalities (e.g., blind or low-vision users, or users with hearing impairments).
    • Adding audio descriptions (AD) and closed captions (CC) is an essential method to enhance video accessibility, but creating these elements is time-consuming and challenging for non-professional authors, making it difficult to accurately identify accessibility issues in videos.
    • Current tools often rely on "speech gaps" to identify accessibility issues, but in many scenarios (e.g., tutorial videos, vlogs, lectures), even without significant speech gaps, critical visual content may remain undescribed or auditory information may not be adequately expressed.
  • Importance:
    • Video accessibility is crucial for equitable information access for blind, low-vision, and hearing-impaired individuals.
    • Enhancing content creators' efficiency and the quality of accessible content while avoiding the omission of key elements.
  • Research Motivation and Related Work:
    • Existing tools can only identify certain accessibility issues, such as silent segments or errors in automatic speech recognition, leading to the omission of critical accessibility problems.
    • There is a need to develop new technologies and systems to more efficiently identify accessibility issues while assisting authors in real-time editing and previewing accessible content.

Solution

  • Proposed Method or Solution:
    • Develop a novel system—CrossA11y—that leverages cross-modal technology to automatically detect accessibility issues based on asymmetry between visual and audio tracks (modal asymmetry).
    • CrossA11y includes:
      • Analysis of video and audio tracks to detect asymmetry.
      • A unified interface allowing authors to view issues, immediately create AD or CC, and preview edited results.
  • Innovations:
    • Application of cross-modal technology to analyze the matching degree between visual and auditory content using machine learning, improving the accuracy of accessibility issue detection.
    • Introduction of a weighting mechanism (time-based weighting) to reduce irrelevant matches between visual and audio segments.
    • Provision of real-time editing tools for visual and audio accessibility information, enhancing efficiency compared to existing tools.
  • Implementation Steps and Key Technologies:
    • Segmentation: Divide the visual track into continuous frame units and the audio track into speech and non-speech units.
    • Cross-modal Alignment: Use machine learning algorithms (e.g., MIL-NCE and MMV models) to calculate matching scores between visual and audio/text segments.
    • Post-processing Filtering: Remove non-critical accessibility issues (e.g., silent background segments and direct-to-camera speech by presenters).
    • Unified Interface Design: Provide color-coded markers on the video timeline to help users quickly navigate and locate issues.

Research Outcomes

  • Specific Results:
    • The CrossA11y system significantly improved authors' efficiency in detecting and resolving video accessibility issues.
    • The developed cross-modal scoring algorithm achieved high recall rates (identification of critical accessibility issues) and moderate precision rates (reducing false positives).
  • Comparative Advantages Over Existing Solutions:
    • Compared to traditional methods relying solely on "speech gaps," CrossA11y demonstrates higher accuracy in detection and cross-modal analysis, identifying accessibility issues in cases where speech and visual content coexist without description.
    • Provides an intuitive user interface that reduces cognitive load for creators while enabling parallel editing with high automation.
  • Experimental or Evaluation Results:
    • Tested system performance using sample videos (20 videos), achieving a recall rate of 98.4% for visual accessibility issues and a precision rate of 98.3% for audio accessibility issues.
    • Experimental users (12 participants) reported reduced task completion time and lower mental workload when using CrossA11y compared to traditional tools.
    • Feedback from two YouTube content creators indicated that CrossA11y effectively saved time in video accessibility processing and improved overall video quality.
  • Limitations and Future Directions:
    • The segmentation algorithm may struggle to accurately distinguish semantic units in certain cases (e.g., videos with a single camera angle).
    • The modal alignment algorithm requires improvement in detecting background visual details or rare visual content in videos.
    • Future efforts will focus on integrating with existing video editing tools to provide end-to-end accessibility support for video production workflows.
    • Explore broader applications of accessibility diagnostics across modalities, such as text and images, GIFs and articles, etc.

In summary, the CrossA11y system provides critical support for identifying and resolving video accessibility issues through efficient cross-modal analysis and intuitive visualization tools, laying a solid foundation for the next generation of accessibility technologies.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/uist/85006/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3526113.3545703
At a Glance

Paper Snapshot

fact_check
dataset
Source
UIST
calendar_month
Year
2022
emoji_events
Award
Best Paper
group
Authors
5 authors
sell
Subtopics
Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille), Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration), Universal & Inclusive Design
work
Professions
Content Creators (YouTubers, Podcasters), UI/UX Designers, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
1 related papers