Utilizing a Dense Video Captioning Technique for Generating Image Descriptions of Comics for People with Visual Impairments
Authors
To improve the accessibility of visual figures, auto-generation of text description of individual images has been studied. However, it cannot be directly applied to comics as the descriptions can be redundant as similar scenes appear in a row. To address this issue, we propose generating the descriptions per group of related images and demonstrate how an dense captioning technique for videos can be utilized for this purpose and ways to improve its performance. To assess the effectiveness of our approach and to identify factors affecting the quality of text descriptions of comics, we conducted a preliminary study with 3 sighted evaluators and a main user study with 12 participants with visual impairments. The results show that text descriptions generated per group of images are perceived to be better than those generated per image in terms of accuracy, clarity, understandability, length, informativeness and preference for sighted groups, when annotator is human. In the same conditions, when the annotator is AI, it exhibited better performance in terms of length. Also, people with visual impairments prefer group descriptions because of conciseness, smooth connectivity of sentences, and non-repetitive features. Based on the findings, we provide design recommendations for generating accessible comic descriptions at a scale for blind users.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 75%
A11yBoard: Making Digital Artboards Accessible to Blind and Low-Vision Users
CHI '23· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +1
- 67%
Understanding Blind Screen-Reader Users' Experiences of Digital Artboards
CHI '21· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 67%
Ga11y: an Automated GIF Annotation System for Visually Impaired Users
CHI '22· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 67%
What makes web data tables accessible? Insights and a tool for rendering accessible tables for people with visual impairments
CHI '22· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 67%
Modeling Touch-based Menu Selection Performance of Blind Users via Reinforcement Learning
CHI '23· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 67%
"It Brought Me Joy": Opportunities for Spatial Browsing in Desktop Screen Readers
CHI '25· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 67%
InSupport: Proxy Interface for Enabling Efficient Non-Visual Interaction with Web Data Records
IUI '22· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 67%
Space-Mag: An Automatic, Scalable, and Rapid Space Compactor for Optimizing Smartphone App Interfaces for Low-Vision Users
UbiComp '23· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +1
- 67%
OmniScribe: Authoring Immersive Audio Descriptions for 360° Videos
UIST '22· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 67%
MagnePins: A Modular, Affordable, and DIY Refreshable Braille and Tactile Display
UIST '25· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
Based on Jaccard similarity of research subtopics & professions (≥60%)