Say It My Way: Exploring Control in Conversational Visual Question Answering with Blind Users

Voice AccessibilityGenerative AI (Text, Image, Music, Video)Human-LLM CollaborationExplainable AI (XAI)Speech-Language Pathologists & AudiologistsHCI Researchers

Paper Title

Say It My Way: Exploring Control in Conversational Visual Question Answering with Blind Users

Publication Info

  • Topic area: Assistive technology for blind users leveraging conversational visual question answering systems.
  • Keywords: Visual question answering, blind users, customization, prompting techniques, assistive AI, generative AI, user control, accessibility, personalization, multimodal systems.

Background and Problem

  • Problem / challenge: Current assistive VQA systems for blind users follow rigid interaction patterns, limiting opportunities for customization and user control over system responses.
  • Significance: Blind users rely on these systems for access to visual information, where efficiency, clarity, and alignment with user goals are critical. Misaligned or verbose responses can lead to confusion, delays, or safety risks.
  • Motivation and related work: Previous research has focused on automated systems for VQA, but little attention has been given to enabling blind users to shape or control these systems. General-purpose generative AI has explored techniques like prompt engineering and fine-tuning, but these approaches remain underexplored in assistive contexts.

Solution

  • Proposed approach: Investigate customization techniques for blind users interacting with a conversational VQA system (Be My AI) through a three-phase study: in-lab sessions, diary studies, and post-study interviews.
  • Novelty:
    1. Empirical findings on how blind users adapt to VQA constraints and use customization techniques.
    2. Introduction of the Ask&Prompt dataset, capturing multi-turn interactions, images, and annotations from real-world and lab contexts.
    3. Identification of design opportunities for greater user control in assistive VQA systems.
  • Procedure and key techniques:
    • Introduced customization techniques include binary feedback, zero-shot style and intention prompting, chain-of-thought prompting, and image-as-text prompting.
    • Conducted a 10-day diary study documenting 363 interactions, followed by post-study interviews to refine annotations and gather deeper insights.
    • Analyzed user behaviors, preferences, and reflections across diverse contexts.

Results

  • Concrete findings:
    • User inputs averaged 10–20 words, while system responses were over ten times longer (median = 160–242 words).
    • Conversations often extended to multiple turns (median = 3, up to 21), especially for iterative tasks like image capture or clarification.
    • Participants actively experimented with customization techniques to manage verbosity, focus on task-relevant details, and strengthen trust.
  • Advantage over baselines:
    • Customization techniques reduced verbosity, improved alignment with user goals, and supported iterative problem-solving compared to default system behavior.
  • Experiments / evaluation:
    • Study included 11 blind participants (aged 25–57) with varying levels of blindness and familiarity with AI.
    • Interactions were analyzed for word counts, conversational turns, success rates, and customization techniques.
    • Contextual factors like environment familiarity, time pressure, and social settings were annotated and analyzed.
  • Limitations and future work:
    • The system lacked persistent customization, limiting user influence across sessions.
    • Study duration was restricted to 10 days to mitigate fatigue from repetitive preferences.
    • Future work could explore memory-enabled systems for persistent user-driven personalization and multimodal customization.

Summary

This study investigates how blind users exert control over conversational VQA systems using customization techniques like binary feedback, style prompting, intention prompting, and chain-of-thought prompting. Through a three-phase study, participants actively shaped responses to manage verbosity, focus on task-relevant details, and strengthen trust. Findings highlight the need for on-demand verbosity control, user-centered spatial references, proactive camera guidance, and verification mechanisms for high-stakes tasks. The Ask&Prompt dataset, capturing multi-turn interactions and annotations, is publicly released to support further research. These contributions aim to advance assistive AI systems that enhance autonomy and personalization for blind users.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222265/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791834
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Voice Accessibility, Generative AI (Text, Image, Music, Video), Human-LLM Collaboration, Explainable AI (XAI)
work
Professions
Speech-Language Pathologists & Audiologists, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers