DanmuA11y: Making Time-Synced On-Screen Video Comments (Danmu) Accessible to Blind and Low Vision Users via Multi-Viewer Audio Discussions

Honorable Mention
Voice AccessibilityAccessible GamingUniversal & Inclusive DesignDisability Service ProvidersAssistive Technology Specialists

Research Background and Issues

What problems or challenges did the authors identify?

  • Barriers in visual-centric design: The video commenting feature "Danmu" (bullet comments) is designed based on visual presentation, making it difficult for visually impaired users to access and understand the content of Danmu.
  • Specific challenges: Through preliminary research, the authors identified three main barriers: lack of visual context, interference between Danmu and video audio, and the disorganized nature of Danmu.

Why is this issue important?

  • Danmu is a popular form of social interaction on video platforms, offering users a shared viewing experience. However, the current design excludes visually impaired users from participating in this interaction, leading to social isolation.
  • This issue is tied to the broader topic of video content accessibility, significantly impacting the viewing experience of many visually impaired users.

Research Motivation and Related Work

  • How the authors understand the issue: By conducting in-depth interviews and collaborative viewing experiments with visually impaired video users, the authors gained a comprehensive understanding of their needs.
  • Related work: While existing research has focused on accessibility design for visual content such as images and short videos, there remains a significant research gap in Danmu accessibility and the design of multi-user social experiences.

Solutions

What methods or solutions did the authors propose?

  • System design: The authors proposed the "DanmuA11y" system, which transforms Danmu into multi-user audio discussions to improve the Danmu experience for visually impaired users.
  • Three core functionalities:
    1. Enhancing the visual context of Danmu.
    2. Seamlessly integrating Danmu into videos to avoid audio interference.
    3. Simulating a shared viewing experience through multi-user audio.

What are the innovative aspects of this solution?

  • Visual-to-audio transformation: By incorporating AI-generated visual descriptions, the system combines video content with Danmu information, helping visually impaired users understand visual elements.
  • Virtual audience design: The system organizes Danmu into simulated dialogues among multiple viewers and presents them through spatial audio to create a sense of social co-viewing.
  • Personalized interaction: Instant Danmu access is enabled through smartphone vibration gestures, making interaction simple and convenient.

What are the implementation steps and key technologies used?

  1. Video segmentation: Identify appropriate insertion points for Danmu based on non-speech and speech segments.
  2. Visual description generation: Use GPT-4 and scene detection technology to provide visual descriptions of keyframes for each video segment.
  3. Danmu content prioritization: Filter high-quality Danmu based on metrics such as informativeness, creativity, and diversity of opinions.
  4. Spatial audio configuration: Position AI narrators and virtual audience members around the video in different locations to enhance audio distinguishability.
  5. User testing and feedback: Validate the system's technical performance and user satisfaction through user studies.

Research Outcomes

What specific outcomes were achieved?

  • Significant improvement in Danmu comprehension: Compared to traditional methods, users demonstrated better understanding of Danmu, with a notable reduction in misunderstandings in video summaries.
  • Enhanced viewing experience: Users reported significantly reduced interference, and the integration of Danmu with video content felt more natural.
  • Strengthened sense of social connection: Users felt closer to other viewers and enjoyed the video-watching process more when using the system.

What advantages does it have over existing solutions?

  • Information augmentation: By parsing visual elements and Danmu content through visual descriptions, visually impaired users can engage more deeply in video discussions.
  • User control: The system allows users to access Danmu at any time via vibration gestures, avoiding the complexity of screen reader operations.
  • Recreating social experiences: By simulating multi-user dialogues and using spatial audio, the system recreates the "co-viewing" experience for visually impaired users.

What were the experimental or evaluation results?

  • User feedback: In a comparative experiment with 12 visually impaired users, DanmuA11y received significantly higher overall ratings than the baseline system in terms of usability, Danmu comprehension, video experience, and sense of social connection.
  • Technical performance: The accuracy of visual description generation reached 92.5%, and the multi-user audio presentations in most scenarios were considered highly realistic and relevant.
  • User habits: Participants activated 87.4% of Danmu notifications on average and exhibited different preferences across various video types.

Limitations and Future Directions

  • Limitations in visual complexity: Rapidly changing scenes (e.g., fast action) pose challenges for the accuracy of visual descriptions.
  • Need for personalized design: Users expressed individualized needs for comment filters, audio settings, and Danmu selection criteria.
  • Long-term impact studies: Further research is needed to explore how DanmuA11y enhances the social capabilities of visually impaired users over time.
  • Integration with real-time comments: Future work could explore applying DanmuA11y to real-time video stream comments.

Through this study, the authors highlight the accessibility challenges of Danmu and demonstrate a technology-driven solution with DanmuA11y. This not only advances accessible design for video-based social interactions but also provides valuable insights for accessibility technologies in other domains.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188308/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713496
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
Honorable Mention
group
Authors
4 authors
sell
Subtopics
Voice Accessibility, Accessible Gaming, Universal & Inclusive Design
work
Professions
Disability Service Providers, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
2 related papers