SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers

Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)Disability Service ProvidersAssistive Technology Specialists

Title of the Paper

SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers

Paper Information

  • Field: Accessibility Technology and Human-Computer Interaction
  • Keywords: Interactive Video Exploration, Audio Description, Accessibility Technology, Low-Vision Adaptation, Video Content Understanding, Spatial Audio, Machine Learning, User Experience, Multimodal Interaction, Video Assistance Tools

Research Background and Problem

  • Problems and Challenges:
    • Current audio descriptions (AD) for videos are static, leading to incomplete information and failing to meet the diverse needs of visually impaired users.
    • Blind or Low-Vision (BLV) users often experience higher cognitive load when relying on AD and lack the ability to independently explore video content compared to sighted audiences.
    • Existing ADs tend to be concise, omitting secondary visual elements or details, which limits user immersion.
  • Significance:
    • With the rapid growth of online video content, especially user-generated content, providing high-quality, personalized description experiences for visually impaired users is crucial for information equity and social inclusion.
  • Motivation and Related Work:
    • Previous research has extensively explored the automation and improvement of video AD generation, but these models often fall short in accuracy, richness of description, and interactivity.
    • Against this backdrop, there is an urgent need for a new approach that leverages intelligent technologies to enhance the granularity and interactive experience of AD.

Solution

  • Method Overview:
    • Propose an AI-supported system called SPICA, which allows BLV users to interact with video content through augmented audio descriptions (including spatial audio and object-level descriptions).
    • SPICA offers timeline navigation features, enabling exploration of keyframes in temporal and spatial dimensions, along with detailed descriptions of key objects in the video.
  • Innovations:
    • Introduces a hierarchical description method, allowing users to delve deeper into scene or object details based on their needs.
    • Utilizes machine learning to automatically generate frame-level and object-level audio descriptions without requiring additional human annotations.
    • Enhances immersion and spatial awareness by incorporating spatial audio and object location descriptions.
    • Provides multiple interaction methods (touch and keyboard operations) to accommodate user preferences.
  • Implementation Steps and Technologies:
    1. Key Modules:
      • Machine Learning Pipeline: Analyzes and segments scenes, detects keyframes and objects, generates object descriptions, and adds spatial audio.
      • Frontend Interaction Features:
        • Video player supporting frame and object exploration.
        • Frame-level description navigation list.
        • Object-level description navigation list.
    2. Technical Tools:
      • Scene Analysis: Uses a Mask RCNN-based model for object segmentation.
      • Description Generation: Employs GPT-4 to optimize natural language descriptions of objects.
      • Spatial Audio Extension: Retrieves object-related sounds from the freesound.org database and spatializes audio using 3D positioning.
      • User Interface: Builds an accessible interactive interface using React and Flask.

Research Outcomes

  • Specific Results:
    • SPICA improves BLV users' understanding and immersion in video content.
    • Users can actively explore video content to access more details, meeting personalized needs.
    • Technical validation demonstrates SPICA's high accuracy in object annotation and description generation, with the quality of automatically generated descriptions surpassing baseline systems.
  • Comparison with Existing Solutions and Advantages:
    • Compared to traditional static AD, SPICA's hierarchical descriptions and interactive features offer a more detailed and flexible experience.
    • Spatial audio and object-level descriptions provide users with a stronger sense of scene presence.
    • Experiments show significant improvements in user-reported content immersion and overall experience scores after using SPICA.
  • Experiments and Evaluation Results:
    • In the experiment, 14 BLV users participated in functionality evaluations of SPICA.
    • Users reported that SPICA was easy to operate (average score of 6.29 out of 7) and significantly enhanced the richness (6.57) and immersion (6.29) of video content.
    • The timeline navigation feature was perceived to improve narrative coherence, though the automatic pause mechanism sparked some debate.
    • The object exploration feature assisted users in spatial understanding, with nearly half (39.4%) of paused frames being further explored.
  • Limitations and Future Directions:
    • Applicability to long videos requires further evaluation.
    • Better personalization is needed, such as adapting to different user preferences for narrative style and description detail.
    • Improve the accuracy of the machine learning module and its ability to perceive user contexts.
    • Explore more efficient interaction methods, including voice interaction and synchronized viewing features for group audiences.

This study provides important practical insights into video accessibility design. Future work will focus on optimizing description generation models, enhancing multimodal interaction experiences, and expanding SPICA's real-world impact through online testing.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147481/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642632
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
8 authors
sell
Subtopics
Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
work
Professions
Disability Service Providers, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
10 related papers