GenAssist: Making Image Generation Accessible

Best Paper
Generative AI (Text, Image, Music, Video)Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)Assistive Technology SpecialistsHCI Researchers

Title of the Paper

GenAssist: Making Image Generation Accessible

Paper Information

  • Subject Area: Accessibility, Generative AI, Visual Creativity Tools
  • Keywords: Accessibility, Image Generation, Generative AI, Blind and Low Vision Users, Creative Support Tools

Research Background and Problem

  • Issues and Challenges:

    • Blind and Low Vision (BLV) creators face significant challenges in effectively generating or retrieving images to meet their visual creation and expression needs.
    • Current image generation tools lack accessibility support for BLV creators, particularly in understanding and evaluating the content and quality of generated images.
  • Significance:

    • Images are a vital medium for communication and expression, yet existing image generation tools are not user-friendly for BLV creators, limiting their creative potential.
    • While text-to-image models can generate high-fidelity images from textual descriptions, BLV creators struggle to assess whether the generated results align with their creative intentions.
  • Research Motivation and Related Work:

    • Previous research has attempted to improve the accessibility of creative tools through tactile devices, audio notifications, or textual descriptions, but little has been done to make image generation tools more accessible.
    • Automatic image description tools exist, but they are primarily designed for image consumption rather than addressing creators' needs for detailed information about generated images, such as style, lighting, or emotion.

Solution

  • Proposed Method:

    • The GenAssist system aims to enhance the experience of BLV creators using text-to-image generation tools by assisting with prompts, image descriptions, and comparisons.
    • The system offers the following functionalities:
      • Verifying alignment with text prompts.
      • Providing details about image elements not specified in the prompts.
      • Summarizing and comparing similarities and differences among generated images.
  • Innovations:

    • Leveraging a large language model (GPT-4) to generate vision-related questions and combining it with vision-language models (BLIP-2 and CLIP) to answer these questions, extracting image content and style information.
    • Presenting information in a tabular format for flexible navigation by screen reader users and generating detailed image descriptions on demand.
  • Implementation Steps and Techniques:

    • Prompt Verification: GPT-4 generates prompt-based verification questions, and BLIP-2 provides answers.
    • Content and Style Extraction: Extracts image content such as setting, subject, and emotion, as well as style elements like medium, lighting, and perspective.
    • Description Summarization: Integrates visual information to generate detailed descriptions for each image and compares similarities and differences across multiple images.
    • Interactive Interface: A screen reader-compatible web interface is developed using Gradio.

Research Outcomes

  • Specific Results and Experimental Findings:

    • GenAssist significantly improved BLV creators' ability to understand and compare images in user studies:
      • Users rated GenAssist as more effective in understanding differences between images, while reducing cognitive load and frustration during tasks.
      • Compared to baseline interfaces, users reported higher satisfaction, confidence, and efficiency in selecting generated images.
    • The system's automatically generated descriptions and answers achieved high coverage, with accuracy exceeding 90% in most visual information categories.
  • Advantages and Features:

    • GenAssist provides a structured approach to help BLV creators quickly comprehend generated images and iteratively refine text prompts.
    • The system enables BLV creators to access critical visual information without needing to ask repetitive questions.
  • Limitations and Future Directions:

    • Limitations:
      • The model occasionally generates incorrect information or omits details, such as errors in classifying image styles, perspectives, or object detection.
      • Limited support for complex charts or infographics.
    • Future Directions:
      • Enhance recognition of subtle differences through richer visual questions and improved models.
      • Develop features supporting information-rich graphics and dynamic content (e.g., videos or animations).
      • Provide more intuitive prompt design guidance, such as visual style recommendations or multimodal input support.

Conclusion

GenAssist offers an innovative solution that enables blind and low vision creators to access and produce high-quality images. Its design explores the potential of combining text-to-image generation technology with visual descriptions. The research findings validate its effectiveness across various tasks, marking a significant advancement in generative AI for accessibility and paving the way for future developments in creative tools.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/uist/126670/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3586183.3606735
At a Glance

Paper Snapshot

fact_check
dataset
Source
UIST
calendar_month
Year
2023
emoji_events
Award
Best Paper
group
Authors
3 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
work
Professions
Assistive Technology Specialists, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers