Slide Gestalt: Automatic Structure Extraction in Slide Decks for Non-Visual Access

Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)Speech-Language Pathologists & AudiologistsAssistive Technology Specialists

Document Title

Slide Gestalt: Automatic Structure Extraction in Slide Decks for Non-Visual Access

Document Information

  • Domain: Human-Computer Interaction and Accessibility Technology
  • Keywords: Accessibility, Slide Structure Extraction, Hierarchical Information, Screen Readers, Multimodal Correspondence and Alignment

Research Background and Problem

  • Identified Issues or Challenges:

    • Current slide presentation software (e.g., PowerPoint or Google Slides) exhibits significant limitations in enabling blind and visually impaired (BVI) users to access content, such as the lack of metadata (e.g., alternative text for images) and incorrect reading order of elements.
    • The visual design of slide content, especially repetitive structures (e.g., titles, separators, and step-by-step slides), performs poorly with screen readers, leading to a time-consuming and confusing reading experience.
    • Existing authoring practices tend to focus on visual structuring but lack non-visual support for these visual intentions.
  • Significance:

    • Slides are a widely used communication medium, but their accessibility directly impacts the efficiency of information acquisition for visually impaired users.
    • Understanding the hierarchical structure and content intent of slides is crucial for efficient navigation by screen reader users.
  • Motivation and Related Work:

    • Previous research has focused on improving the accessibility of basic slide elements (e.g., adding alternative text), but there remains a significant gap in interpreting the overall structure and style of slides.
    • Literature analysis has identified common slide structures, such as separator slides, step-by-step slides, and topic-split slides, which provide important research clues for navigating slides for visually impaired users.

Solution

  • Proposed Method or Solution:

    • The authors propose an automated processing system, "Slide Gestalt," to identify the hierarchical structure of slide decks and support visually impaired users in efficiently navigating content.
    • By computing visual and textual correspondences between slides, the system generates a multi-level hierarchy.
  • Innovations:

    • Utilizes multimodal data, including textual and visual features, combined with neural embedding models and sequence alignment algorithms to extract structured information from slides.
    • Interface design incorporates a two-level structure display: a high-level overview of sections and a fine-grained description of slide groups, providing users with flexible navigation options.
  • Implementation Steps and Key Techniques:

    1. Data Extraction: Use the Google Slides API to obtain slide elements, including text, images, and layouts.
    2. Correspondence Calculation: Generate an inter-slide distance matrix based on visual and textual embeddings and perform clustering using a sequence alignment algorithm.
    3. Structure Detection: Detect separator slides, step-by-step slides, and topic-split slides, while distinguishing standalone slides.
    4. Hierarchy Generation and Description: Extract titles, key elements, and ranges of slide groups, and adjust the reading order to optimize the screen reader experience.

Research Outcomes

  • Specific Results:

    • Slide Gestalt successfully identified and generated hierarchical structures of slides, with experimental validation showing its effectiveness (F1 score = 0.81).
    • User studies demonstrated that participants using the system navigated faster, encountered less redundant information, and achieved higher accuracy in slide browsing and information retrieval tasks.
  • Advantages Over Existing Solutions:

    • Compared to the linear reading mode of traditional screen readers, Slide Gestalt is more effective in helping visually impaired users understand content structure and reduce information redundancy.
    • The system's ability to process multimodal data enhances the overall comprehension of slide content.
  • Experimental or Evaluation Results:

    • In user evaluations, participants using the hierarchical interface significantly reduced the time taken to answer questions (45.8 seconds vs. 96.9 seconds) and the number of unnecessary elements read (6.5 vs. 20.6).
    • All participants found the hierarchical interface more helpful for understanding content and supporting more flexible reading and navigation.
  • Limitations and Future Directions:

    • One limitation is the system's inability to detect the impact of animations on content, which may result in occlusion issues.
    • The current method requires direct access to slide elements, making it challenging to handle restricted formats like PDFs.
    • Improvements are needed in recursive structure inference for hierarchy generation, with future work considering multimodal reasoning based on large language models.
    • Exploration of hybrid tools to enable slide authors to embed structural information during the early stages of creation, laying the groundwork for subsequent improvements.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/96522/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3580921
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
work
Professions
Speech-Language Pathologists & Audiologists, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
10 related papers