Say It My Way: Exploring Control in Conversational Visual Question Answering with Blind Users
Authors
Paper Title
Say It My Way: Exploring Control in Conversational Visual Question Answering with Blind Users
Publication Info
- Topic area: Assistive technology for blind users leveraging conversational visual question answering systems.
- Keywords: Visual question answering, blind users, customization, prompting techniques, assistive AI, generative AI, user control, accessibility, personalization, multimodal systems.
Background and Problem
- Problem / challenge: Current assistive VQA systems for blind users follow rigid interaction patterns, limiting opportunities for customization and user control over system responses.
- Significance: Blind users rely on these systems for access to visual information, where efficiency, clarity, and alignment with user goals are critical. Misaligned or verbose responses can lead to confusion, delays, or safety risks.
- Motivation and related work: Previous research has focused on automated systems for VQA, but little attention has been given to enabling blind users to shape or control these systems. General-purpose generative AI has explored techniques like prompt engineering and fine-tuning, but these approaches remain underexplored in assistive contexts.
Solution
- Proposed approach: Investigate customization techniques for blind users interacting with a conversational VQA system (Be My AI) through a three-phase study: in-lab sessions, diary studies, and post-study interviews.
- Novelty:
- Empirical findings on how blind users adapt to VQA constraints and use customization techniques.
- Introduction of the Ask&Prompt dataset, capturing multi-turn interactions, images, and annotations from real-world and lab contexts.
- Identification of design opportunities for greater user control in assistive VQA systems.
- Procedure and key techniques:
- Introduced customization techniques include binary feedback, zero-shot style and intention prompting, chain-of-thought prompting, and image-as-text prompting.
- Conducted a 10-day diary study documenting 363 interactions, followed by post-study interviews to refine annotations and gather deeper insights.
- Analyzed user behaviors, preferences, and reflections across diverse contexts.
Results
- Concrete findings:
- User inputs averaged 10–20 words, while system responses were over ten times longer (median = 160–242 words).
- Conversations often extended to multiple turns (median = 3, up to 21), especially for iterative tasks like image capture or clarification.
- Participants actively experimented with customization techniques to manage verbosity, focus on task-relevant details, and strengthen trust.
- Advantage over baselines:
- Customization techniques reduced verbosity, improved alignment with user goals, and supported iterative problem-solving compared to default system behavior.
- Experiments / evaluation:
- Study included 11 blind participants (aged 25–57) with varying levels of blindness and familiarity with AI.
- Interactions were analyzed for word counts, conversational turns, success rates, and customization techniques.
- Contextual factors like environment familiarity, time pressure, and social settings were annotated and analyzed.
- Limitations and future work:
- The system lacked persistent customization, limiting user influence across sessions.
- Study duration was restricted to 10 days to mitigate fatigue from repetitive preferences.
- Future work could explore memory-enabled systems for persistent user-driven personalization and multimodal customization.
Summary
This study investigates how blind users exert control over conversational VQA systems using customization techniques like binary feedback, style prompting, intention prompting, and chain-of-thought prompting. Through a three-phase study, participants actively shaped responses to manage verbosity, focus on task-relevant details, and strengthen trust. Findings highlight the need for on-demand verbosity control, user-centered spatial references, proactive camera guidance, and verification mechanisms for high-stakes tasks. The Ask&Prompt dataset, capturing multi-turn interactions and annotations, is publicly released to support further research. These contributions aim to advance assistive AI systems that enhance autonomy and personalization for blind users.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)