Supporting Novices Author Audio Descriptions via Automatic Feedback
Authors
Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)Universal & Inclusive DesignGame AccessibilityAssistive Technology SpecialistsFreelancers (Design, Writing, Translation)
Document Title
Supporting Novices Author Audio Descriptions via Automatic Feedback
Document Information
- Subject Areas: Human-Computer Interaction, Video Assistive Technology, Natural Language Processing
- Keywords: Accessibility Technology, Blind Assistance, Audio Description, Automatic Feedback, Scene Description, Computer Vision, Natural Language Processing
Research Background and Issues
-
Problems or Challenges:
- Audio Description (AD) is a crucial accessibility technology designed to help blind or visually impaired individuals understand video content. However, most videos lack audio descriptions—for instance, only about 0.004% of videos on Amazon Prime Video include audio descriptions.
- Creating audio descriptions requires professional expertise, which is costly and time-consuming, with costs ranging from $12 to $75 per minute and delivery times potentially extending up to a week.
- Attempts to automate audio description generation have so far failed to capture critical information about character actions and positions, which are essential for visually impaired users.
-
Importance:
- As video becomes one of the primary mediums for information dissemination, the lack of audio descriptions limits the ability of visually impaired individuals to fully engage with such content.
- Reducing the cost and time required to produce audio descriptions while maintaining quality can promote the broader adoption of accessibility technologies and enhance social inclusivity.
-
Research Motivation and Related Work:
- Previous studies have explored automated and semi-automated methods to reduce the cost of generating audio descriptions, but these methods have not fully addressed the issue of low-quality descriptions.
- Collaborative tools for Scene Description (SD) writing have been shown to improve description quality for novices, and integrating automatic feedback mechanisms could further enhance efficiency.
Solution
-
Method or Solution:
- Developed an audio description creation tool that integrates manual scene description with real-time automated feedback. The tool leverages computer vision and natural language processing technologies to generate feedback suggestions, assisting users in refining scene descriptions.
- Real-time feedback is presented in the form of interactive word clouds, allowing authors to select representative tags to include in their scene descriptions. The tool aims to optimize both the content and process of descriptions, ensuring efficiency and quality.
-
Innovations:
- Proposed a novel automatic feedback mechanism combining computer vision (using Amazon Rekognition) and natural language processing.
- Used word clouds to visually guide novices in identifying potentially missing critical visual elements.
- Optimized tag selection algorithms to reduce redundancy and irrelevant annotations.
-
Implementation Steps and Key Technologies:
- Utilize Amazon Rekognition for scene understanding and generate initial tags.
- Optimize tags through text matching and clustering algorithms (e.g., DBSCAN) to produce representative tags.
- Perform real-time semantic matching between tags and text as users compose scene descriptions, providing reference suggestions.
- Design a user interface that visualizes the scene description creation process and allows real-time modifications.
Research Outcomes
-
Specific Results:
- Completed the design of a system supporting scene description creation with real-time automatic feedback functionality and conducted a controlled experiment with 60 participants.
- The automatic feedback tool improved the descriptiveness, objectivity, and learning quality of scene descriptions.
- In terms of cost reduction, the automatic feedback system lowered production costs by 45%.
-
Advantages Compared to Existing Solutions:
- Automatic feedback significantly reduced the time required for manual feedback evaluation, demonstrating potential as a replacement for fully manual processes.
- Compared to scenarios without feedback, automatic feedback improved description quality while emphasizing objectivity as a key dimension.
-
Experimental Results and Evaluation:
- Experiments showed that human feedback outperformed automatic feedback in terms of richness and inspiration, but automatic feedback had a clear advantage in saving production time.
- Observations in practice indicated that automatic feedback descriptions were more concise, with consistent grammar and vocabulary.
-
Limitations and Future Directions:
- The automatic feedback mechanism is limited by the output quality of current computer vision technologies (e.g., Amazon Rekognition). In certain scenarios, insufficient description information or low-quality tags posed challenges.
- Future exploration could focus on hybrid feedback approaches (combining human and automatic feedback) and real-time dynamic optimization systems tailored to user needs.
- Developing a prior mechanism for video type classification could help select appropriate feedback mechanisms based on task complexity.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How can automated feedback help novices write high-quality audio descriptions?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
- How can real-time automated feedback mechanisms improve efficiency and quality of scene description?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
- How can feedback mechanisms integrating computer vision and NLP support audio description creation?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
lightbulb
Practical Problems
1- Blind people struggle to access audio descriptions of video content; production is costly and time-consuming.Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
- 67%
Diffscriber: Describing Visual Design Changes to Support Mixed-Ability Collaborative Presentation Authoring
UIST '22· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +1
- 60%
Screen Recognition: Creating Accessibility Metadata for Mobile Applications from Pixels
CHI '21· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581023
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille), Universal & Inclusive Design, Game Accessibility
work
Professions
Assistive Technology Specialists, Freelancers (Design, Writing, Translation)
article
Content Status
Full text indexed
hub
Related Papers
2 related papers