From Provenance to Aberrations: Image Creator and Screen Reader User Perspectives on Alt Text for AI-Generated Images

Generative AI (Text, Image, Music, Video)Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)Universal & Inclusive DesignConsumers & ShoppersAssistive Technology SpecialistsHCI Researchers

Document Title

From Image Provenance to Aberrations: Alternative Text for AI-Generated Images from the Perspective of Blind Users

Document Information

  • Research Area: Accessibility Computing, AI-Generated Content, Blind User Experience
  • Keywords: AI Art, Text-to-Image Generation, Accessibility, Alternative Text, Blind Users, Screen Reader Users

Research Background and Issues

  • Identified Problems or Challenges:
    1. During the rapid proliferation of AI-generated images, most of these images lack alternative text tailored for screen reader users (SRUs), resulting in insufficient accessibility.
    2. Even in scenarios where alternative text exists, its quality is often inconsistent, with some mainstream web content entirely missing alternative text (e.g., 30% of images lack alternative text).
    3. There is currently no systematic study on alternative text strategies for AI-generated images, nor clear descriptive standards to address this new type of media.
  • Significance: The unique characteristics of AI-generated images (e.g., style combinations, illusions, and unknown provenance information) introduce new dimensions for accessible image descriptions. Properly describing the accessibility of AI images is crucial, especially in preventing misleading information.
  • Research Motivation and Related Work:
    • Globally, over 15 billion AI-generated images exist, yet these images have extremely low accessibility for blind users who rely on textual descriptions.
    • Previous accessibility research has explored alternative text needs for traditional images and complex information (e.g., scientific paper visualizations, social media), but the needs for AI-generated images remain underexplored.
    • With the widespread adoption of generative models (e.g., DALL-E, Midjourney) and the expansion of application scenarios, studying the impact of these contents on screen reader users to make them more perceivable and understandable is increasingly necessary.

Solution

  • Proposed Methods and Solutions: The authors designed a comprehensive approach to systematically study alternative text from various sources, including text prompts used during image generation, alternative text written by creators, alternative text written by professional accessibility experts, and automatically generated text description models.
  • Innovations:
    1. Proposed a framework to evaluate the quality of alternative text from different sources.
    2. Introduced AI-specific descriptive requirements, such as the necessity of provenance information, descriptions of image generation aberrations, and novel style and media expressions.
    3. Collected and analyzed diverse participant preferences (image creators and screen reader users) regarding alternative text quality and content.
  • Implementation Steps and Key Techniques:
    1. Participant Evaluation: Recruited 16 AI image creators and 16 frequent screen reader users for the experiment.
    2. Alternative Text Encoding and Classification:
      • Created four types of alternative text: text prompts used during image generation, text written by creators, text written by experts, and text generated by automated alternative text models.
      • Evaluated each image's alternative text for length, certain linguistic features (nouns, verbs, adjectives, etc.), and lexical uniqueness.
    3. Combining Subjective and Objective Evaluation: Participants scored and compared the four types of alternative text and wrote ideal versions of descriptions.
    4. Quantitative and Qualitative Analysis: Used statistical methods to understand linguistic features and text preference trends, while conducting thematic analysis of participant experiences and suggestions.

Research Outcomes

  • Specific Findings:
    1. Produced a complete dataset of 64 AI-generated images and their four types of alternative text.
    2. Quantitative analysis of texts from different sources revealed:
      • Expert-written texts were more detailed and ranked highest but were sometimes overly verbose.
      • Screen reader users preferred concise and specifically descriptive content in their ideal versions.
      • Text prompts used during generation often contained technical jargon and were rarely rated as ideal descriptions.
    3. Using generation prompts as alternative text has limited informational potential, especially when prompts include irrelevant or unrendered content not present in the image.
    4. Identified key descriptive needs: considering image provenance, generation aberrations, media and style.
  • Comparative Advantages of Existing Solutions:
    • Compared to solely relying on model-generated text or traditional descriptive methods, the authors’ approach comprehensively considers image context, user preferences, and technical feasibility, demonstrating better targeting and readability in descriptions.
  • Experimental Results: The method received positive feedback from participants on both "description comprehensiveness" and "description adaptability." Additionally, based on statistical analysis of participants’ ideal versions, the authors proposed a more dynamic "layered description structure" for alternative text strategies.
  • Limitations and Future Directions:
    • Limitations:
      1. The study primarily focused on broad general scenarios and did not delve into specific contexts (e.g., differences between social media and academic publishing).
      2. Alternative text generated by models (V2L) may not represent the capabilities of the latest large-scale multimodal models.
    • Future Directions:
      1. Explore differences in alternative text needs across various application scenarios.
      2. Further evaluate and enhance descriptive accuracy using new perception and reasoning models (e.g., LLMs).
      3. Investigate better methods for embedding image provenance information into generated text, especially in cases of watermark absence or image tampering.

Output Format

The article provides an in-depth exploration of the core accessibility elements conveyed by AI images for blind users, offering a solid reference framework for industries to develop accessibility strategies for AI-generated images.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/148306/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642325
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille), Universal & Inclusive Design
work
Professions
Consumers & Shoppers, Assistive Technology Specialists, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
7 related papers