VideoA11y: Method and Dataset for Accessible Video Description
Authors
Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)Public Transit OperatorsPhysicians, Nurses & CliniciansAssistive Technology Specialists
Research Background and Issues
- Identified Problems and Challenges: Current AI models perform poorly in generating video descriptions for blind or low-vision (BLV) users due to the suboptimal quality of training data annotations. This results in descriptions that fail to fully meet the needs of BLV users. Additionally, human-generated descriptions are often incomplete or contain grammatical, spelling, or semantic errors.
- Significance: The rapid growth of video content has further widened the information gap for BLV individuals. While audio descriptions (AD) can address these needs, their production requires expensive professional teams. This is particularly problematic for user-generated content (e.g., YouTube and TikTok), where audio descriptions are scarce and lag behind demand.
- Research Motivation and Related Work: Previous studies have primarily focused on assisting human describers in generating descriptions, while fully automated approaches have yet to effectively address the challenge of producing detailed video descriptions that adhere to AD guidelines. Furthermore, the limitations of existing datasets hinder the quality and usability of AI-generated descriptions.
Solution
- Proposed Method and Innovations:
- A novel approach, VideoA11y, is proposed, which combines multimodal large language models (MLLMs) with video accessibility guidelines to design standardized prompts for generating highly clear and accurate video descriptions.
- Based on this method, a VideoA11y-40K dataset containing 40,000 video clips has been constructed, making it the largest and most comprehensive video description dataset tailored for BLV users.
- Implementation Steps:
- Extract and organize 42 AD guidelines from professional audio description resources to guide description generation.
- Use MLLMs such as GPT-4 Vision (GPT-4V) to generate video descriptions, integrating keyframe data from videos.
- Construct and classify the VideoA11y-40K dataset, categorizing videos into 15 classes, with descriptions generated using standardized prompts for each video.
- Key Technologies Used:
- Multimodal language models (e.g., GPT-4V and Video-LLaVA) for zero-shot prompt optimization.
- Video keyframe extraction algorithms (based on local maxima algorithms) to ensure significant changes in videos are captured.
- Dataset preparation, including video classification and rigorous quality evaluation.
Research Outcomes
- Specific Outcomes:
- Descriptions generated by VideoA11y outperform novice human annotations in clarity, accuracy, objectivity, and descriptiveness, approaching the quality of high-quality human annotations.
- The constructed VideoA11y-40K dataset significantly improves the quality of descriptions generated by multimodal language models, providing a robust data foundation for video accessibility.
- A new benchmark standard is proposed for evaluating video description models designed for BLV users.
- Advantages:
- VideoA11y-generated descriptions surpass novice human annotations in terms of accuracy and clarity, closely matching professional human annotations.
- The VideoA11y method significantly enhances existing models while reducing reliance on human involvement, offering high scalability.
- Reduced hallucination phenomena (i.e., generating inaccurate content), with notable improvements in accuracy when human annotations are used as references.
- Experiments and Evaluation Results:
- Across five user studies (involving 347 sighted users, 40 BLV users, and 7 professional describers), VideoA11y demonstrated outstanding performance:
- BLV users preferred VideoA11y-generated descriptions in over 90% of scenarios.
- Professional describers favored VideoA11y-generated descriptions, considering them aligned with professional standards.
- Technical experiments showed that open-source MLLM models fine-tuned with the VideoA11y-40K dataset significantly outperformed baseline models without fine-tuning.
- Across five user studies (involving 347 sighted users, 40 BLV users, and 7 professional describers), VideoA11y demonstrated outstanding performance:
- Limitations and Future Directions:
- Limitations:
- In the absence of human annotations, minor hallucination phenomena may occur (e.g., introducing details not present in the video).
- The current method cannot fully meet BLV users' personalized needs (e.g., preferences for description length or level of detail).
- Seamless integration of descriptions with video content (e.g., inserting natural pause points) has not been fully achieved.
- Future Directions:
- Optimize generation models to reduce hallucination scenarios, for example, by employing direct preference optimization (DPO) techniques or integrating auxiliary models to enhance input accuracy.
- Collect BLV user preference data to enable dynamic adjustment of description style and level of detail.
- Integrate with existing systems (e.g., Rescribe) to support embedded descriptions and seamless presentation within video streams.
- Limitations:
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How can multimodal large language models combined with video accessibility guidelines improve clarity and accuracy of AI-generated video descriptions for blind or low vision users?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
- Can building a large dataset of high-quality video descriptions (e.g., VideoA11y-40K) significantly improve performance of existing multimodal language models?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
- Can AI-generated video descriptions better meet professional audio description standards than those produced by beginners?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
lightbulb
Practical Problems
1- Blind users struggle to efficiently access video content information, and existing audio descriptions are scarce and costly.Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
- 75%
Towards More Accessible Scientific PDFs for People with Visual Impairments: Step-by-Step PDF Remediation to Improve Tag Accuracy
CHI '25· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 60%
ALAP: Accessible LaTeX Based Mathematical Document Authoring and Presentation
CHI '19· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +1
- 60%
VIPBoard: Improving Screen-Reader Keyboard for Visually Impaired People with Character-Level Auto Correction
CHI '19· Voice Accessibility +1
- 60%
Infosonics: Accessible Infographics for People who are Blind using Sonification and Voice
CHI '22· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +1
- 60%
Toucha11y: Making Inaccessible Public Touchscreens Accessible
CHI '23· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +1
- 60%
Accessibility of Profile Pictures: Alt Text and Beyond to Express Identity Online
CHI '23· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +1
- 60%
ChitChatGuide: Conversational Interaction Using Large Language Models for Assisting People with Visual Impairments to Explore a Shopping Mall
MobileHCI '24· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714096
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
work
Professions
Public Transit Operators, Physicians, Nurses & Clinicians, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
7 related papers