How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision People

Honorable Mention
Human-LLM CollaborationVisual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)Voice AccessibilityPsychiatrists & PsychotherapistsSpeech-Language Pathologists & AudiologistsAssistive Technology Specialists

Paper Title

How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision People

Publication Info

  • Topic area: Accessibility and AI-powered visual interpretation for Blind and Low Vision (BLV) users.
  • Keywords: Multimodal large language models, visual interpretation, Blind and Low Vision, accessibility, conversational AI, diary study, user satisfaction, trust, GPT-4o, visual assistant.

Background and Problem

  • Problem / challenge: Existing AI-powered visual interpretation tools often produce inaccurate or incomplete descriptions, fail to meet BLV users' daily needs reliably, and lack conversational capabilities for goal-directed assistance.
  • Significance: Reliable visual interpretation systems can enhance BLV individuals' independence in daily tasks like cooking, navigating, and managing medications, reducing reliance on sighted assistance.
  • Motivation and related work: Prior systems like Seeing AI and Orcam have shown limited accuracy and usability in real-world scenarios. Recent advances in multimodal large language models (MLLMs) promise improved descriptive accuracy and conversational interaction, but their real-world performance and impact on BLV users remain underexplored.

Solution

  • Proposed approach: VisionPal, an MLLM-enabled visual interpretation application, was developed to study how BLV users interact with such systems in real-world contexts over a two-week diary study.
  • Novelty:
    1. Empirical insights into real-world usage patterns, goals, and challenges of MLLM-enabled visual interpretation systems by BLV users.
    2. Evidence of MLLM strengths (e.g., high descriptive accuracy) and limitations (e.g., conversational errors and hallucinations).
    3. Definition of the "visual assistant" skill and nine associated behaviors to guide future system design.
  • Procedure and key techniques:
    • Conducted a two-week diary study with 20 BLV participants using VisionPal.
    • Collected 551 diary entries, including photos, AI-generated descriptions, follow-up conversations, and user feedback.
    • Analyzed accuracy, satisfaction, trust, and conversational interactions.
    • Proposed design principles and behavioral guidelines for MLLM-based visual assistants.

Results

  • Concrete findings:
    • Initial photo descriptions were highly accurate (mean accuracy = 2.91/3; 91.8% with no hallucinations).
    • Satisfaction (mean = 4.13/5) and trust (mean = 3.76/5) ratings were generally high.
    • Conversational follow-up responses were correct in only 56.6% of cases, with 22.2% containing false information.
    • Text-based queries had the highest error rate (34.6% incorrect or partially incorrect responses).
  • Advantage over baselines:
    • VisionPal achieved higher satisfaction, trust, and descriptive accuracy compared to pre-MLLM systems (e.g., prior studies reported satisfaction = 2.76/5 and trust = 2.43/4).
  • Experiments / evaluation:
    • Participants used VisionPal across diverse settings (e.g., homes, work, transit) and for various goals (e.g., identifying objects, reading text, understanding scenes).
    • 375 conversations were analyzed, revealing patterns of engagement and challenges in follow-up interactions.
  • Limitations and future work:
    • Observer effect and participants’ expertise with visual interpretation systems may have influenced results.
    • Limited representation of novice users and data collection during the holiday season may have introduced biases.
    • Future work should address conversational accuracy gaps and incorporate more diverse user perspectives.

Summary

This study investigated how MLLM-enabled visual interpretation applications support BLV users in real-world contexts. VisionPal demonstrated high accuracy in generating descriptive visual interpretations, leading to increased user satisfaction and trust compared to pre-MLLM systems. However, conversational interactions revealed significant accuracy gaps, particularly for text-based queries and follow-up questions. The paper defines the "visual assistant" skill, proposing nine behaviors to guide the development of more effective and user-aligned visual interpretation systems. Future efforts should focus on improving conversational accuracy, contextual adaptation, and personalization to better meet BLV users' needs.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222234/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3793266
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
Honorable Mention
group
Authors
5 authors
sell
Subtopics
Human-LLM Collaboration, Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille), Voice Accessibility
work
Professions
Psychiatrists & Psychotherapists, Speech-Language Pathologists & Audiologists, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
0 related papers