"It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language Models

Honorable Mention
Explainable AI (XAI)Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)Universal & Inclusive DesignVoice AccessibilitySpeech-Language Pathologists & AudiologistsCommunity Health WorkersAssistive Technology Specialists

Paper Title

'It's trained by non-disabled people': Evaluating How Image Quality Affects Product Captioning with Vision-Language Models

Publication Info

  • Topic area: Vision-Language Models (VLMs) for accessibility in image captioning.
  • Keywords: Vision-Language Models, image quality, blind and low-vision users, product identification, accessibility, image captioning, human-computer interaction, dataset annotation, model evaluation, assistive technology.

Background and Problem

  • Problem / challenge: Vision-Language Models (VLMs) struggle to accurately caption images with common quality issues (e.g., blur, framing, rotation) encountered by blind and low-vision (BLV) users. Existing evaluations often overlook these real-world challenges.
  • Significance: Accurate and detailed captions are critical for BLV users to safely and effectively identify products, especially for tasks involving health or safety, such as identifying food or medication.
  • Motivation and related work: Prior research has focused on high-quality images and general captioning metrics, neglecting the impact of real-world image quality issues on VLM performance. This study addresses the gap by systematically evaluating VLMs on BLV-relevant datasets with degraded images.

Solution

  • Proposed approach: A two-part study: (1) a survey of BLV users to understand their experiences with VLM-based tools, and (2) a systematic evaluation of VLMs on a curated dataset of product images with varying quality issues.
  • Novelty:
    1. Development of a dataset of 1,859 BLV-captured product images annotated with product type, brand, and variety.
    2. Systematic evaluation of four VLMs (GPT-4.1, Gemini 2.5 Flash, Llama 3.2 90B, Molmo 72B) on high- and low-quality images.
    3. Analysis of how specific image quality issues (blur, framing, rotation) and product properties (e.g., rounded labels) affect VLM performance.
    4. Recommendations for improving VLM reliability and user support.
  • Procedure and key techniques:
    1. Conducted a survey with 86 BLV participants to gather insights on VLM usage and challenges.
    2. Curated and annotated a dataset of high- and low-quality product images from the VizWiz dataset.
    3. Evaluated VLMs using a structured annotation scheme and logistic regression to analyze performance across image quality issues and product properties.
    4. Proposed improvements in data curation, training objectives, and inference-time techniques for VLMs.

Results

  • Concrete findings:
    • VLM accuracy drops significantly for low-quality images, with GPT achieving 74.9% accuracy (down from 98.5% on high-quality images), Gemini 71.7%, Llama 44.1%, and Molmo 36.1%.
    • Blur reduces correct identification odds by 88.3%, framing by 84.5%, and rotation by 79.5%.
    • Products with both rounded labels and text panels further challenge Gemini, Llama, and Molmo, reducing accuracy to 67.5%, 42.5%, and 30.0%, respectively.
  • Advantage over baselines:
    • GPT and Gemini outperform open-source models (Llama, Molmo) on both high- and low-quality images.
    • GPT shows better resilience to image quality issues compared to other models.
  • Experiments / evaluation:
    • Dataset: 729 high-quality and 1,130 low-quality images annotated for product type, brand, and variety.
    • Metrics: Accuracy of product identification across image quality issues and product properties.
    • Models: GPT-4.1, Gemini 2.5 Flash, Llama 3.2 90B, Molmo 72B.
  • Limitations and future work:
    • Focused on U.S.-based products and English-speaking users; cross-cultural and multilingual evaluations are needed.
    • Binary treatment of image quality issues; future work should quantify degradation levels.
    • Did not evaluate holistic caption quality (e.g., hedging language, uncertainty communication).

Summary

This study highlights the challenges VLMs face in accurately captioning low-quality images taken by BLV users, with significant performance drops for common issues like blur, framing, and rotation. A curated dataset and systematic evaluation of four VLMs reveal that while GPT and Gemini perform better, open-source models like Llama and Molmo lag behind. The findings underscore the need for disability-centered model evaluation, improved datasets, and targeted training and inference techniques to enhance VLM reliability. These insights pave the way for developing more accessible and robust AI tools for BLV users.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/221844/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791309
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
Honorable Mention
group
Authors
7 authors
sell
Subtopics
Explainable AI (XAI), Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille), Universal & Inclusive Design, Voice Accessibility
work
Professions
Speech-Language Pathologists & Audiologists, Community Health Workers, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
0 related papers