Beyond Visual Perception: Insights from Smartphone Interaction of Visually Impaired Users with Large Multimodal Models

Human-LLM CollaborationVisual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)Disability Service ProvidersAssistive Technology Specialists

Research Background and Issues

  • What problems or challenges did the authors identify?
    The paper explores the capabilities and limitations of large multimodal models (LMMs, such as GPT-4) in providing visual assistance to users with visual impairments. While these emerging technologies can help users better understand their environment through natural language descriptions, the authors identified two major issues:

    • The system's limited contextual awareness, including "hallucinations" (i.e., generating nonexistent details) and misunderstandings of social scenes, styles, and human identities.
    • Inability to correctly interpret or execute user goals based on user intent, resulting in insufficient operational support.
  • Why is this issue important?
    Users with visual impairments require effective assistive technologies to perform daily tasks. As AI technology advances, exploring its limitations and development potential in real-world applications is crucial for improving technology design and enhancing user experience.

  • Research Motivation and Related Work
    Existing studies have focused on using generative AI platforms for visual assistance, such as describing objects in images and answering visual questions. However, these studies primarily emphasize technical performance or content acquisition, lacking in-depth exploration of user interaction methods in real-life scenarios. This study aims to depict how LMM technology impacts the daily lives and social interactions of visually impaired individuals by introducing user cases and experiences.


Solutions

  • What methods or solutions did the authors propose?
    The authors proposed strategies to overcome the limitations of these technologies, including improving interactions between users and AI tools and introducing human assistance when necessary. Additionally, they suggested addressing task execution issues through multi-agent systems (human-human and human-AI collaboration) and future AI-AI cooperation.

  • What is innovative about the solution?

    • Emphasized dynamic handoff mechanisms among users, AI tools, and remote sighted assistants (RSAs).
    • Proposed an AI "Deferral Learning" framework for handling sensitive content and identity recognition tasks.
    • Explored potential applications of real-time video analysis to address data insufficiency and navigation risks in static image processing.
  • What are the implementation steps and key technologies used?

    • Conducted semi-structured interviews to collect user experiences with BMA (Be My AI).
    • Summarized system performance and limitations from interviews, supplemented with image description data from social media.
    • Analyzed user experiences such as "real-time feedback," "goal understanding," and "goal support," and proposed design recommendations based on long-term and short-term memory AI mechanisms.
    • Designed collaborative interactions across users, AI assistants, and human assistants to optimize workflow.

Research Outcomes

  • What specific outcomes were achieved?

    • BMA enhances spatial awareness, social interaction understanding, and object recognition capabilities for visually impaired users through visual descriptions.
    • The system frequently exhibits "hallucinations" when describing visual scenes, such as adding nonexistent details or misinterpreting emotions and identities of humans and animals.
    • Demonstrates critical support deficiencies in task execution, such as lacking clear follow-up guidance or real-time feedback.
  • How does it compare to existing solutions?

    • LMM models offer more natural language interaction capabilities, surpassing traditional visual assistance applications like "Seeing AI."
    • Proposed improvements in coordination among users, AI tools, and RSAs, enhancing applicability in social and task-based scenarios.
  • What were the experimental or evaluation results?
    Interviews revealed that users could achieve greater independence in specific tasks using BMA, but limitations required the introduction of RSAs or reliance on users' spatial memory and auditory perception to solve problems, such as locating dropped objects or verifying AI hallucinations.

  • Limitations and Future Directions

    • Limitations:
      • BMA lacks conversation history storage and contextual memory functionality.
      • Overemphasis on static images, with limited real-time feedback capabilities, restricting navigation task efficiency.
      • Descriptions of emotions and identities are overly subjective, sometimes leading to controversy or inaccuracies.
    • Future Directions:
      • Develop real-time video processing technology to improve interaction efficiency.
      • Introduce automatic switching mechanisms between human and AI collaboration to enhance complex task completion.
      • Expand research scope to broader cultural contexts and multilingual environments to validate the technology's universality.

Conclusion

By analyzing the experiences of visually impaired users with BMA, this study reveals the practical application potential and shortcomings of LMM technology. The research provides guidance for designing smarter, more interactive, and personalized assistive technologies in the future, proposing specific improvement directions such as enhancing handoff mechanisms, introducing real-time video analysis, and enabling AI-AI collaboration. These findings are significant for advancing technology development in the field of visual assistance.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189592/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714210
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Human-LLM Collaboration, Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
work
Professions
Disability Service Providers, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
10 related papers