Supporting Novices Author Audio Descriptions via Automatic Feedback

Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)Universal & Inclusive DesignGame AccessibilityAssistive Technology SpecialistsFreelancers (Design, Writing, Translation)

Document Title

Supporting Novices Author Audio Descriptions via Automatic Feedback

Document Information

  • Subject Areas: Human-Computer Interaction, Video Assistive Technology, Natural Language Processing
  • Keywords: Accessibility Technology, Blind Assistance, Audio Description, Automatic Feedback, Scene Description, Computer Vision, Natural Language Processing

Research Background and Issues

  • Problems or Challenges:

    • Audio Description (AD) is a crucial accessibility technology designed to help blind or visually impaired individuals understand video content. However, most videos lack audio descriptions—for instance, only about 0.004% of videos on Amazon Prime Video include audio descriptions.
    • Creating audio descriptions requires professional expertise, which is costly and time-consuming, with costs ranging from $12 to $75 per minute and delivery times potentially extending up to a week.
    • Attempts to automate audio description generation have so far failed to capture critical information about character actions and positions, which are essential for visually impaired users.
  • Importance:

    • As video becomes one of the primary mediums for information dissemination, the lack of audio descriptions limits the ability of visually impaired individuals to fully engage with such content.
    • Reducing the cost and time required to produce audio descriptions while maintaining quality can promote the broader adoption of accessibility technologies and enhance social inclusivity.
  • Research Motivation and Related Work:

    • Previous studies have explored automated and semi-automated methods to reduce the cost of generating audio descriptions, but these methods have not fully addressed the issue of low-quality descriptions.
    • Collaborative tools for Scene Description (SD) writing have been shown to improve description quality for novices, and integrating automatic feedback mechanisms could further enhance efficiency.

Solution

  • Method or Solution:

    • Developed an audio description creation tool that integrates manual scene description with real-time automated feedback. The tool leverages computer vision and natural language processing technologies to generate feedback suggestions, assisting users in refining scene descriptions.
    • Real-time feedback is presented in the form of interactive word clouds, allowing authors to select representative tags to include in their scene descriptions. The tool aims to optimize both the content and process of descriptions, ensuring efficiency and quality.
  • Innovations:

    • Proposed a novel automatic feedback mechanism combining computer vision (using Amazon Rekognition) and natural language processing.
    • Used word clouds to visually guide novices in identifying potentially missing critical visual elements.
    • Optimized tag selection algorithms to reduce redundancy and irrelevant annotations.
  • Implementation Steps and Key Technologies:

    1. Utilize Amazon Rekognition for scene understanding and generate initial tags.
    2. Optimize tags through text matching and clustering algorithms (e.g., DBSCAN) to produce representative tags.
    3. Perform real-time semantic matching between tags and text as users compose scene descriptions, providing reference suggestions.
    4. Design a user interface that visualizes the scene description creation process and allows real-time modifications.

Research Outcomes

  • Specific Results:

    • Completed the design of a system supporting scene description creation with real-time automatic feedback functionality and conducted a controlled experiment with 60 participants.
    • The automatic feedback tool improved the descriptiveness, objectivity, and learning quality of scene descriptions.
    • In terms of cost reduction, the automatic feedback system lowered production costs by 45%.
  • Advantages Compared to Existing Solutions:

    • Automatic feedback significantly reduced the time required for manual feedback evaluation, demonstrating potential as a replacement for fully manual processes.
    • Compared to scenarios without feedback, automatic feedback improved description quality while emphasizing objectivity as a key dimension.
  • Experimental Results and Evaluation:

    • Experiments showed that human feedback outperformed automatic feedback in terms of richness and inspiration, but automatic feedback had a clear advantage in saving production time.
    • Observations in practice indicated that automatic feedback descriptions were more concise, with consistent grammar and vocabulary.
  • Limitations and Future Directions:

    • The automatic feedback mechanism is limited by the output quality of current computer vision technologies (e.g., Amazon Rekognition). In certain scenarios, insufficient description information or low-quality tags posed challenges.
    • Future exploration could focus on hybrid feedback approaches (combining human and automatic feedback) and real-time dynamic optimization systems tailored to user needs.
    • Developing a prior mechanism for video type classification could help select appropriate feedback mechanisms based on task complexity.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/96052/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581023
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille), Universal & Inclusive Design, Game Accessibility
work
Professions
Assistive Technology Specialists, Freelancers (Design, Writing, Translation)
article
Content Status
Full text indexed
hub
Related Papers
2 related papers