TableNarrator: Making Image Tables Accessible to Blind and Low Vision People
Authors
Research Background and Problem
-
What issues or challenges did the authors identify?
The authors pointed out that the widespread use of image-based tables poses significant challenges to information access for blind and low-vision (BLV) individuals. Traditional screen readers and optical character recognition (OCR) technologies are unable to effectively handle the complex structure and semantic information of image-based tables, making it difficult for BLV users to fully comprehend or interact with these tables. -
Why is this problem important?
Tables, as one of the most important forms of data presentation, are widely used in fields such as scientific research, education, and business analysis. The lack of accessibility for image-based tables limits BLV users' participation in these critical areas, thereby impacting equity and social inclusion. -
Research Motivation and Related Work
Although artificial intelligence technologies (e.g., table recognition and table understanding) have made some progress, they still have limitations in meeting the needs of BLV users. For instance, they struggle with complex table formats and often produce outputs with "redundant" or "hallucinatory" (inaccurate) information. This study aims to address these issues by investigating the needs of BLV users and designing a system that meets those needs.
Solution
-
What methods or solutions did the authors propose?
The authors developed a system called TableNarrator to improve the accessibility of image-based tables. This system leverages artificial intelligence technologies to extract the structure and semantic information of tables and provides BLV users with customizable alternative text through a concise and clear interaction model. -
What are the innovative aspects of this solution?
- The system is designed based on the needs of BLV users, including in-depth exploration of table metadata (e.g., number of rows, columns, headers) and content semantics.
- It offers a straightforward interaction model, enabling BLV users to access image-based tables as flexibly as they would actual tables.
- Users can customize the granularity and order of information in the alternative text according to personal preferences.
- For generating alternative text, the system combines table layout analysis and structural recognition with large language models (LLMs) to produce comprehensible table content summaries.
-
What are the implementation steps and key technologies used?
TableNarrator consists of three main modules:- Table Layout Analysis Module: Detects and segments table regions in images, including headers, footers, and body content, to facilitate subsequent tasks.
- Table Structure Recognition Module: Extracts the position of each cell, row-column relationships, and complex semantic relationships (e.g., key-value pairs), using multimodal algorithms and LLMs to generate accurate alternative text.
- Personalization and Interaction Module: Provides users with flexible interaction experiences through gestures (e.g., single-finger or two-finger swipes) and customization options, supporting access to information at various levels of granularity.
Research Outcomes
-
What specific outcomes were achieved?
- Technical Evaluation: Experiments demonstrated that TableNarrator outperformed baseline methods, including GPT-4V, in table metadata accuracy, content extraction, and semantic analysis.
- User Study: Users highly appreciated TableNarrator's interaction model, with the system achieving a usability score of 90.6 (excellent range) and showing lower cognitive load and higher efficiency in workload tests.
-
What advantages does it have compared to existing solutions?
- Better adaptation to complex table structures and more accurate information delivery compared to baseline methods like GPT-4V and VoiceOver.
- Reduced redundant information and improved browsing efficiency through customization and interaction optimization.
- A user-friendly interface that allows users to adjust the granularity and order of output text based on their needs, enhancing flexibility.
-
What were the experimental or evaluation results?
- In technical evaluations, TableNarrator demonstrated high accuracy in tasks such as table detection, logical cell positioning, and cell relationship classification, surpassing existing domain models in related metrics.
- In user evaluations, TableNarrator was described as "very intuitive" and "highly efficient," with its customizable options being considered more targeted and reducing users' learning costs.
-
Limitations and Future Directions
- TableNarrator's performance on complex layouts or tables in photographic images still needs improvement. Multi-round interaction and error correction mechanisms could further enhance the system's robustness.
- Its ability to handle multilingual tables needs to be strengthened to serve BLV users from diverse linguistic backgrounds and avoid biases caused by language limitations.
- More advanced information processing features, such as QA dialogue modules and data analysis functionalities, could be introduced to provide users with more creative interaction options.
Conclusion
Through modular design and innovative interaction models, TableNarrator significantly enhances the accessibility of image-based tables for BLV users. Its outstanding performance in technical and user evaluations validates its effectiveness and potential. By deeply integrating user needs, the system surpasses existing solutions in terms of information delivery, interaction convenience, and customization. However, parsing complex tables and ensuring multilingual compatibility remain key areas for future research and optimization. This work not only provides an important new tool for the BLV community but also sets a benchmark for research on the accessibility of image-based tables.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can accessibility of graphical tables be improved for blind and low vision (BLV) users?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
- How can interaction models enable BLV users to flexibly access table information?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
- How can complex structure and semantic information in graphical tables be accurately extracted and converted into understandable alt text?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
Practical Problems
1- BLV users struggle to understand information in graphical tables.Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
- 100%
Towards Enabling Blind People to Independently Write on Printed Forms
CHI '19· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 100%
Examining Visual Semantic Understanding in Blind and Low-Vision Technology Users
CHI '21· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 100%
Scene Text Access: A Comparison of Mobile OCR Modalities for Blind Users
IUI '19· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 67%
A Face Recognition Application for People with Visual Impairments: Understanding Use Beyond the Lab
CHI '18· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 67%
Enabling People with Visual Impairments to Navigate Virtual Reality with a Haptic and Auditory Cane Simulation
CHI '18· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 67%
SteeringWheel: A Locality-Preserving Magnification Interface for Low Vision Web Browsing
CHI '18· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 67%
Caption Crawler: Enabling Reusable Alternative Text Descriptions using Reverse Image Search
CHI '18· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 67%
Accessible Maps for the Blind: Comparing 3D Printed Models with Tactile Graphics
CHI '18· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +1
- 67%
Exploring the Opportunities for Technologies to Enhance Quality of Life with People who have Experienced Vision Loss
CHI '19· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 67%
Virtual Reality Without Vision: A Haptic and Auditory White Cane to Navigate Complex Virtual Worlds
CHI '20· Social & Collaborative VR +1
Based on Jaccard similarity of research subtopics & professions (≥60%)