GenAssist: Making Image Generation Accessible
Best PaperAuthors
Title of the Paper
GenAssist: Making Image Generation Accessible
Paper Information
- Subject Area: Accessibility, Generative AI, Visual Creativity Tools
- Keywords: Accessibility, Image Generation, Generative AI, Blind and Low Vision Users, Creative Support Tools
Research Background and Problem
-
Issues and Challenges:
- Blind and Low Vision (BLV) creators face significant challenges in effectively generating or retrieving images to meet their visual creation and expression needs.
- Current image generation tools lack accessibility support for BLV creators, particularly in understanding and evaluating the content and quality of generated images.
-
Significance:
- Images are a vital medium for communication and expression, yet existing image generation tools are not user-friendly for BLV creators, limiting their creative potential.
- While text-to-image models can generate high-fidelity images from textual descriptions, BLV creators struggle to assess whether the generated results align with their creative intentions.
-
Research Motivation and Related Work:
- Previous research has attempted to improve the accessibility of creative tools through tactile devices, audio notifications, or textual descriptions, but little has been done to make image generation tools more accessible.
- Automatic image description tools exist, but they are primarily designed for image consumption rather than addressing creators' needs for detailed information about generated images, such as style, lighting, or emotion.
Solution
-
Proposed Method:
- The GenAssist system aims to enhance the experience of BLV creators using text-to-image generation tools by assisting with prompts, image descriptions, and comparisons.
- The system offers the following functionalities:
- Verifying alignment with text prompts.
- Providing details about image elements not specified in the prompts.
- Summarizing and comparing similarities and differences among generated images.
-
Innovations:
- Leveraging a large language model (GPT-4) to generate vision-related questions and combining it with vision-language models (BLIP-2 and CLIP) to answer these questions, extracting image content and style information.
- Presenting information in a tabular format for flexible navigation by screen reader users and generating detailed image descriptions on demand.
-
Implementation Steps and Techniques:
- Prompt Verification: GPT-4 generates prompt-based verification questions, and BLIP-2 provides answers.
- Content and Style Extraction: Extracts image content such as setting, subject, and emotion, as well as style elements like medium, lighting, and perspective.
- Description Summarization: Integrates visual information to generate detailed descriptions for each image and compares similarities and differences across multiple images.
- Interactive Interface: A screen reader-compatible web interface is developed using Gradio.
Research Outcomes
-
Specific Results and Experimental Findings:
- GenAssist significantly improved BLV creators' ability to understand and compare images in user studies:
- Users rated GenAssist as more effective in understanding differences between images, while reducing cognitive load and frustration during tasks.
- Compared to baseline interfaces, users reported higher satisfaction, confidence, and efficiency in selecting generated images.
- The system's automatically generated descriptions and answers achieved high coverage, with accuracy exceeding 90% in most visual information categories.
- GenAssist significantly improved BLV creators' ability to understand and compare images in user studies:
-
Advantages and Features:
- GenAssist provides a structured approach to help BLV creators quickly comprehend generated images and iteratively refine text prompts.
- The system enables BLV creators to access critical visual information without needing to ask repetitive questions.
-
Limitations and Future Directions:
- Limitations:
- The model occasionally generates incorrect information or omits details, such as errors in classifying image styles, perspectives, or object detection.
- Limited support for complex charts or infographics.
- Future Directions:
- Enhance recognition of subtle differences through richer visual questions and improved models.
- Develop features supporting information-rich graphics and dynamic content (e.g., videos or animations).
- Provide more intuitive prompt design guidance, such as visual style recommendations or multimodal input support.
- Limitations:
Conclusion
GenAssist offers an innovative solution that enables blind and low vision creators to access and produce high-quality images. Its design explores the potential of combining text-to-image generation technology with visual descriptions. The research findings validate its effectiveness across various tasks, marking a significant advancement in generative AI for accessibility and paving the way for future developments in creative tools.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can blind and low-vision (BLV) creators more effectively use text-to-image generation tools?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
- How can these intelligent tools help BLV users understand the content and style of generated images?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
- Through what means can BLV creators' satisfaction and confidence during image generation be improved?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
Practical Problems
1- Blind and low-vision users struggle to evaluate whether generated images match their creative intent.Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
- 100%
Social Sense-making with AI: Designing an Open-ended AI experience with a Blind Child
CHI '21· Generative AI (Text, Image, Music, Video) +1
- 100%
The Sky is the Limit: Understanding How Generative AI can Enhance Screen Reader Users' Experience with Productivity Applications
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 75%
Twitter A11y: A Browser Extension to Make Twitter Images Accessible
CHI '20· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 75%
Everyday Uncertainty: How Blind People Use GenAI Tools for Information Access
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 67%
From Provenance to Aberrations: Image Creator and Screen Reader User Perspectives on Alt Text for AI-Generated Images
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 67%
Creating Disability Story Videos with Generative AI: Motivation, Expression, and Sharing
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 60%
Rich Representations of Visual Content for Screen Reader Users
CHI '18· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +1
- 60%
Reading with the Tongue: Individual Differences Affect the Perception of Ambiguous Stimuli with the BrainPort
CHI '20· Vibrotactile Feedback & Skin Stimulation +1
- 60%
Introducing the Gamer Information-Control Framework: Enabling Access to Digital Games for People with Visual Impairment
CHI '20· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +1
- 60%
Improving Colour Patterns to Assist People with Colour Vision Deficiency
CHI '22· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)