Investigating Use Cases of AI-Powered Scene Description Applications for Blind and Low Vision People
Authors
Document Title
Investigating Use Cases of AI-Powered Scene Description Applications for Blind and Low Vision People
Document Information
- Research Area: Human-Computer Interaction, Assistive Technology, Computer Vision
- Keywords: Use Cases, Scene Description, Artificial Intelligence, Blind and Low Vision People, Computer Vision, Assistive Technology, Diary Study
Research Background and Problem
- Identified Issues or Challenges:
- Many existing scene description applications help blind and low vision (BLV) people access visual information, but most studies focus on applications requiring human remote assistants, without adequately exploring the use of AI-based scene description applications.
- BLV people's trust and satisfaction with AI-generated descriptions are relatively low, with an average trust score of 2.43 (out of 4) and satisfaction score of 2.76 (out of 5).
- Significance:
- With rapid advancements in AI technology, AI-based assistive tools have the potential to reduce dependence on human assistants and enhance BLV people's independence. However, their effectiveness and applicability require further investigation.
- Research Motivation and Related Work:
- Comparisons of assistive applications like VizWiz and Aira, which rely on human assistants, show high applicability in daily tasks. However, AI-based scene description applications (e.g., Seeing AI) still have many unknowns regarding their ability to "meet BLV people's needs."
- Preliminary studies indicate that BLV users are not always satisfied with descriptions provided by Seeing AI, with only 43% of descriptions deemed satisfactory and generally low accuracy levels.
Solution
- Proposed Solution:
- Design and develop an AI-based scene description application to observe and analyze BLV people's use cases for such applications.
- Conduct a two-week diary study with 16 BLV participants to record their interactions and experiences with the description tool.
- Innovative Aspects of the Solution:
- Introduced the first systematic study of BLV users' interactions with AI-based scene description applications, examining the interplay between AI-generated descriptions and users' existing knowledge and visual information goals.
- Explored how BLV users creatively use the application (e.g., avoiding hazards, resolving visual disputes) and test AI capabilities to understand its limits.
- Implementation Steps:
- Application Development: Develop the application using Microsoft Azure AI Vision image analysis API to generate photo scene descriptions and record user feedback.
- Experiment Design: Train participants and require them to submit at least one photo daily over two weeks, documenting location, information needs, satisfaction, and trust scores.
- Data Analysis: Combine diary data with follow-up semi-structured interviews, using qualitative coding to identify use cases and user behavior patterns. Employ linear mixed-effects models to analyze factors influencing trust and satisfaction.
Research Outcomes
- Specific Findings:
- Common use cases include:
- Identifying objects and their attributes (e.g., color, shape, etc.);
- Understanding environmental information;
- Testing and learning AI description accuracy.
- Unique user needs:
- Avoiding contact with hazardous or unclean items;
- Determining the presence of people;
- Capturing emotional moments (e.g., pet interactions).
- Scene description performance data:
- Average description accuracy score was 1.95 (out of 3), indicating moderate overall performance.
- Satisfaction and trust in descriptions were significantly correlated with accuracy, though users could sometimes infer useful information from partially correct descriptions.
- Common use cases include:
- Advantages:
- Highlighted the automation and privacy features of AI-based scene description tools, which may outperform human assistants in certain scenarios.
- AI was perceived by users as an "objective" source of information, useful for resolving visual disputes.
- Experiment or Evaluation Results:
- Users' trust and satisfaction with AI varied across different use cases and description accuracy levels.
- Participants preferred using the application in personal spaces (e.g., at home), aligning with their information needs.
- AI's limitations in complex tasks hindered widespread adoption, but its core visual information support remained valuable for decision-making.
- Limitations and Future Directions:
- Limitations:
- Diary studies cannot fully capture long-term usage behaviors.
- Short-term experiments may overestimate the proportion of "testing the application" use cases.
- Future Directions:
- Investigate how integrating more features (e.g., OCR, image exploration) can expand the applicability of AI-based scene description tools.
- Optimize feedback mechanisms between users and AI to improve model performance while allowing users to dynamically adjust trust and satisfaction.
- Further explore potential use cases and feedback mechanisms for scene description applications incorporating generative AI technologies.
- Limitations:
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- What practical information needs of blind and low-vision users can AI-generated scene descriptions satisfy?Category: Blind and Low-Vision AccessibilitySimilar questionsarrow_forward
- What factors affect blind and low-vision users' trust in and satisfaction with AI descriptions?Category: Blind and Low-Vision AccessibilitySimilar questionsarrow_forward
- How do blind and low-vision users creatively use features of AI scene description applications?Category: Blind and Low-Vision AccessibilitySimilar questionsarrow_forward
Practical Problems
1- Blind and low-vision users cannot conveniently access important visual information in their surroundings.Category: Blind and Low-Vision AccessibilitySimilar questionsarrow_forward
- 71%
GesturePod: Enabling On-device Gesture-based Interaction for White Cane Users
UIST '19· Hand Gesture Recognition +3
- 67%
Editing Spatial Layouts through Tactile Templates for People with Visual Impairments
CHI '19· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +1
- 67%
Emoji Accessibility for Visually Impaired People
CHI '20· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +1
- 67%
Smartphone Usage by Expert Blind Users
CHI '21· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +1
- 67%
"I Shake The Package To Check If It's Mine": A Study of Package Fetching Practices and Challenges of Blind and Low Vision People in China
CHI '22· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +1
- 67%
Investigating User Perceptions of Epilepsy-Related Seizure Triggers in Mobile Apps: An Analysis of User Reviews
CHI '25· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +1
- 67%
Enabling Tabular Data Exploration for Blind and Low-Vision Users
DIS '24· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +1
- 67%
TangibleGrid: Tangible Web Layout Design for Blind Users
UIST '22· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +1
- 63%
What Do We Mean by "Accessibility Research"? A Systematic Review of Accessibility Papers in CHI and ASSETS from 1994 to 2019
CHI '21· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +3
Based on Jaccard similarity of research subtopics & professions (≥60%)