"It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language Models
Honorable MentionAuthors
Paper Title
'It's trained by non-disabled people': Evaluating How Image Quality Affects Product Captioning with Vision-Language Models
Publication Info
- Topic area: Vision-Language Models (VLMs) for accessibility in image captioning.
- Keywords: Vision-Language Models, image quality, blind and low-vision users, product identification, accessibility, image captioning, human-computer interaction, dataset annotation, model evaluation, assistive technology.
Background and Problem
- Problem / challenge: Vision-Language Models (VLMs) struggle to accurately caption images with common quality issues (e.g., blur, framing, rotation) encountered by blind and low-vision (BLV) users. Existing evaluations often overlook these real-world challenges.
- Significance: Accurate and detailed captions are critical for BLV users to safely and effectively identify products, especially for tasks involving health or safety, such as identifying food or medication.
- Motivation and related work: Prior research has focused on high-quality images and general captioning metrics, neglecting the impact of real-world image quality issues on VLM performance. This study addresses the gap by systematically evaluating VLMs on BLV-relevant datasets with degraded images.
Solution
- Proposed approach: A two-part study: (1) a survey of BLV users to understand their experiences with VLM-based tools, and (2) a systematic evaluation of VLMs on a curated dataset of product images with varying quality issues.
- Novelty:
- Development of a dataset of 1,859 BLV-captured product images annotated with product type, brand, and variety.
- Systematic evaluation of four VLMs (GPT-4.1, Gemini 2.5 Flash, Llama 3.2 90B, Molmo 72B) on high- and low-quality images.
- Analysis of how specific image quality issues (blur, framing, rotation) and product properties (e.g., rounded labels) affect VLM performance.
- Recommendations for improving VLM reliability and user support.
- Procedure and key techniques:
- Conducted a survey with 86 BLV participants to gather insights on VLM usage and challenges.
- Curated and annotated a dataset of high- and low-quality product images from the VizWiz dataset.
- Evaluated VLMs using a structured annotation scheme and logistic regression to analyze performance across image quality issues and product properties.
- Proposed improvements in data curation, training objectives, and inference-time techniques for VLMs.
Results
- Concrete findings:
- VLM accuracy drops significantly for low-quality images, with GPT achieving 74.9% accuracy (down from 98.5% on high-quality images), Gemini 71.7%, Llama 44.1%, and Molmo 36.1%.
- Blur reduces correct identification odds by 88.3%, framing by 84.5%, and rotation by 79.5%.
- Products with both rounded labels and text panels further challenge Gemini, Llama, and Molmo, reducing accuracy to 67.5%, 42.5%, and 30.0%, respectively.
- Advantage over baselines:
- GPT and Gemini outperform open-source models (Llama, Molmo) on both high- and low-quality images.
- GPT shows better resilience to image quality issues compared to other models.
- Experiments / evaluation:
- Dataset: 729 high-quality and 1,130 low-quality images annotated for product type, brand, and variety.
- Metrics: Accuracy of product identification across image quality issues and product properties.
- Models: GPT-4.1, Gemini 2.5 Flash, Llama 3.2 90B, Molmo 72B.
- Limitations and future work:
- Focused on U.S.-based products and English-speaking users; cross-cultural and multilingual evaluations are needed.
- Binary treatment of image quality issues; future work should quantify degradation levels.
- Did not evaluate holistic caption quality (e.g., hedging language, uncertainty communication).
Summary
This study highlights the challenges VLMs face in accurately captioning low-quality images taken by BLV users, with significant performance drops for common issues like blur, framing, and rotation. A curated dataset and systematic evaluation of four VLMs reveal that while GPT and Gemini perform better, open-source models like Llama and Molmo lag behind. The findings underscore the need for disability-centered model evaluation, improved datasets, and targeted training and inference techniques to enhance VLM reliability. These insights pave the way for developing more accessible and robust AI tools for BLV users.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)