TableNarrator: Making Image Tables Accessible to Blind and Low Vision People

Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)Disability Service Providers

Research Background and Problem

  • What issues or challenges did the authors identify?
    The authors pointed out that the widespread use of image-based tables poses significant challenges to information access for blind and low-vision (BLV) individuals. Traditional screen readers and optical character recognition (OCR) technologies are unable to effectively handle the complex structure and semantic information of image-based tables, making it difficult for BLV users to fully comprehend or interact with these tables.

  • Why is this problem important?
    Tables, as one of the most important forms of data presentation, are widely used in fields such as scientific research, education, and business analysis. The lack of accessibility for image-based tables limits BLV users' participation in these critical areas, thereby impacting equity and social inclusion.

  • Research Motivation and Related Work
    Although artificial intelligence technologies (e.g., table recognition and table understanding) have made some progress, they still have limitations in meeting the needs of BLV users. For instance, they struggle with complex table formats and often produce outputs with "redundant" or "hallucinatory" (inaccurate) information. This study aims to address these issues by investigating the needs of BLV users and designing a system that meets those needs.

Solution

  • What methods or solutions did the authors propose?
    The authors developed a system called TableNarrator to improve the accessibility of image-based tables. This system leverages artificial intelligence technologies to extract the structure and semantic information of tables and provides BLV users with customizable alternative text through a concise and clear interaction model.

  • What are the innovative aspects of this solution?

    • The system is designed based on the needs of BLV users, including in-depth exploration of table metadata (e.g., number of rows, columns, headers) and content semantics.
    • It offers a straightforward interaction model, enabling BLV users to access image-based tables as flexibly as they would actual tables.
    • Users can customize the granularity and order of information in the alternative text according to personal preferences.
    • For generating alternative text, the system combines table layout analysis and structural recognition with large language models (LLMs) to produce comprehensible table content summaries.
  • What are the implementation steps and key technologies used?
    TableNarrator consists of three main modules:

    1. Table Layout Analysis Module: Detects and segments table regions in images, including headers, footers, and body content, to facilitate subsequent tasks.
    2. Table Structure Recognition Module: Extracts the position of each cell, row-column relationships, and complex semantic relationships (e.g., key-value pairs), using multimodal algorithms and LLMs to generate accurate alternative text.
    3. Personalization and Interaction Module: Provides users with flexible interaction experiences through gestures (e.g., single-finger or two-finger swipes) and customization options, supporting access to information at various levels of granularity.

Research Outcomes

  • What specific outcomes were achieved?

    1. Technical Evaluation: Experiments demonstrated that TableNarrator outperformed baseline methods, including GPT-4V, in table metadata accuracy, content extraction, and semantic analysis.
    2. User Study: Users highly appreciated TableNarrator's interaction model, with the system achieving a usability score of 90.6 (excellent range) and showing lower cognitive load and higher efficiency in workload tests.
  • What advantages does it have compared to existing solutions?

    • Better adaptation to complex table structures and more accurate information delivery compared to baseline methods like GPT-4V and VoiceOver.
    • Reduced redundant information and improved browsing efficiency through customization and interaction optimization.
    • A user-friendly interface that allows users to adjust the granularity and order of output text based on their needs, enhancing flexibility.
  • What were the experimental or evaluation results?

    • In technical evaluations, TableNarrator demonstrated high accuracy in tasks such as table detection, logical cell positioning, and cell relationship classification, surpassing existing domain models in related metrics.
    • In user evaluations, TableNarrator was described as "very intuitive" and "highly efficient," with its customizable options being considered more targeted and reducing users' learning costs.
  • Limitations and Future Directions

    • TableNarrator's performance on complex layouts or tables in photographic images still needs improvement. Multi-round interaction and error correction mechanisms could further enhance the system's robustness.
    • Its ability to handle multilingual tables needs to be strengthened to serve BLV users from diverse linguistic backgrounds and avoid biases caused by language limitations.
    • More advanced information processing features, such as QA dialogue modules and data analysis functionalities, could be introduced to provide users with more creative interaction options.

Conclusion

Through modular design and innovative interaction models, TableNarrator significantly enhances the accessibility of image-based tables for BLV users. Its outstanding performance in technical and user evaluations validates its effectiveness and potential. By deeply integrating user needs, the system surpasses existing solutions in terms of information delivery, interaction convenience, and customization. However, parsing complex tables and ensuring multilingual compatibility remain key areas for future research and optimization. This work not only provides an important new tool for the BLV community but also sets a benchmark for research on the accessibility of image-based tables.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188971/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714329
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
9 authors
sell
Subtopics
Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
work
Professions
Disability Service Providers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers