How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision People
Honorable MentionAuthors
Paper Title
How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision People
Publication Info
- Topic area: Accessibility and AI-powered visual interpretation for Blind and Low Vision (BLV) users.
- Keywords: Multimodal large language models, visual interpretation, Blind and Low Vision, accessibility, conversational AI, diary study, user satisfaction, trust, GPT-4o, visual assistant.
Background and Problem
- Problem / challenge: Existing AI-powered visual interpretation tools often produce inaccurate or incomplete descriptions, fail to meet BLV users' daily needs reliably, and lack conversational capabilities for goal-directed assistance.
- Significance: Reliable visual interpretation systems can enhance BLV individuals' independence in daily tasks like cooking, navigating, and managing medications, reducing reliance on sighted assistance.
- Motivation and related work: Prior systems like Seeing AI and Orcam have shown limited accuracy and usability in real-world scenarios. Recent advances in multimodal large language models (MLLMs) promise improved descriptive accuracy and conversational interaction, but their real-world performance and impact on BLV users remain underexplored.
Solution
- Proposed approach: VisionPal, an MLLM-enabled visual interpretation application, was developed to study how BLV users interact with such systems in real-world contexts over a two-week diary study.
- Novelty:
- Empirical insights into real-world usage patterns, goals, and challenges of MLLM-enabled visual interpretation systems by BLV users.
- Evidence of MLLM strengths (e.g., high descriptive accuracy) and limitations (e.g., conversational errors and hallucinations).
- Definition of the "visual assistant" skill and nine associated behaviors to guide future system design.
- Procedure and key techniques:
- Conducted a two-week diary study with 20 BLV participants using VisionPal.
- Collected 551 diary entries, including photos, AI-generated descriptions, follow-up conversations, and user feedback.
- Analyzed accuracy, satisfaction, trust, and conversational interactions.
- Proposed design principles and behavioral guidelines for MLLM-based visual assistants.
Results
- Concrete findings:
- Initial photo descriptions were highly accurate (mean accuracy = 2.91/3; 91.8% with no hallucinations).
- Satisfaction (mean = 4.13/5) and trust (mean = 3.76/5) ratings were generally high.
- Conversational follow-up responses were correct in only 56.6% of cases, with 22.2% containing false information.
- Text-based queries had the highest error rate (34.6% incorrect or partially incorrect responses).
- Advantage over baselines:
- VisionPal achieved higher satisfaction, trust, and descriptive accuracy compared to pre-MLLM systems (e.g., prior studies reported satisfaction = 2.76/5 and trust = 2.43/4).
- Experiments / evaluation:
- Participants used VisionPal across diverse settings (e.g., homes, work, transit) and for various goals (e.g., identifying objects, reading text, understanding scenes).
- 375 conversations were analyzed, revealing patterns of engagement and challenges in follow-up interactions.
- Limitations and future work:
- Observer effect and participants’ expertise with visual interpretation systems may have influenced results.
- Limited representation of novice users and data collection during the holiday season may have introduced biases.
- Future work should address conversational accuracy gaps and incorporate more diverse user perspectives.
Summary
This study investigated how MLLM-enabled visual interpretation applications support BLV users in real-world contexts. VisionPal demonstrated high accuracy in generating descriptive visual interpretations, leading to increased user satisfaction and trust compared to pre-MLLM systems. However, conversational interactions revealed significant accuracy gaps, particularly for text-based queries and follow-up questions. The paper defines the "visual assistant" skill, proposing nine behaviors to guide the development of more effective and user-aligned visual interpretation systems. Future efforts should focus on improving conversational accuracy, contextual adaptation, and personalization to better meet BLV users' needs.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)