(Computer) Vision in Action: Comparing Remote Sighted Assistance and a Multimodal Voice Agent in Inspection Sequences
Authors
Paper Title
(Computer) Vision in Action: Comparing Remote Sighted Assistance and a Multimodal Voice Agent in Inspection Sequences
Publication Info
- Topic area: Comparative analysis of human and AI-powered assistance in assistive tasks for blind individuals.
- Keywords: remote sighted assistance, multimodal voice agents, proactivity, human-AI collaboration, assistive technology, ethnomethodology, conversation analysis, blind and low vision, collaborative practices, turn-taking.
Background and Problem
- Problem / challenge: Current multimodal voice agents lack the ability to initiate and modify actions based on real-time environmental cues, limiting their effectiveness in collaborative tasks compared to human assistants.
- Significance: Addressing this gap is crucial for improving assistive technologies for blind and low-vision (BLV) individuals, who rely on such tools for everyday tasks like cleanliness inspection.
- Motivation and related work: While prior studies have explored the capabilities of multimodal voice agents and remote sighted assistance, they have not provided fine-grained analyses of the collaborative practices that underpin successful human assistance. This paper aims to fill that gap by comparing interactions involving a human assistant and a voice agent.
Solution
- Proposed approach: A comparative ethnomethodological conversation analysis of interactions between a blind participant and two types of assistants: a human remote sighted assistant (RSA) and a multimodal voice agent (ChatGPT multimodal “voice mode”).
- Novelty:
- Detailed identification of proactive practices in human assistance that are absent in AI-powered assistance.
- Empirical specification of the conditions under which assistive collaboration succeeds or fails.
- Ethical discussion of the implications of enabling voice agents to initiate environmentally occasioned actions.
- Procedure and key techniques:
- Data collection involved video-recorded interactions of a blind participant performing a stain-finding task with both a human RSA and a multimodal voice agent.
- Analysis used ethnomethodological conversation analysis (EMCA) to examine turn-taking, mutual adjustments, and task coordination.
- Transcripts were created using Jeffersonian and Mondadian conventions to capture speech and embodied actions.
Results
- Concrete findings:
- The human RSA initiated actions based on visual cues, modified actions in real-time, and distributed perceptual work with the participant, leading to task success.
- The voice agent failed to initiate or adapt actions based on environmental cues, resulting in a rigid, step-by-step interaction structure that did not locate the stain.
- Advantage over baselines:
- Human RSA demonstrated superior proactivity and adaptability, enabling effective collaboration and task completion.
- Voice agents lacked the ability to respond to or initiate actions based on real-time visual data, highlighting a significant limitation.
- Experiments / evaluation:
- Two fragments from larger corpora were analyzed: one involving a human RSA using the Be My Eyes app and another involving ChatGPT multimodal “voice mode”.
- Metrics included the ability to locate the stain, the nature of turn-taking, and the distribution of perceptual work.
- Limitations and future work:
- The study's findings are based on a small sample size and specific configurations of the voice agent, limiting generalizability.
- Future research should explore a broader range of tasks, settings, and assistive technologies, as well as the ethical implications of proactive AI behaviors.
Summary
This study compared the collaborative practices of a human remote sighted assistant and a multimodal voice agent in assisting a blind participant with a stain-finding task. The human assistant's success relied on proactive behaviors, including initiating actions based on visual cues, real-time adjustments, and distributing perceptual work. In contrast, the voice agent's rigid, step-by-step interaction failed to achieve the task. The findings highlight the importance of proactivity and environmentally responsive actions for effective assistive technologies. However, the study also raises ethical concerns about enabling AI to autonomously initiate actions, as this could impose value judgments on users. Future work should address these ethical and technical challenges to improve the design of assistive voice agents.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 75%
WESPER: Zero-shot and Realtime Whisper to Normal Voice Conversion for Whisper-based Speech interactions
CHI '23· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 60%
"Nobody Speaks that Fast!" An Empirical Study of Speech Rate in Conversational Agents for People with Vision Impairments
CHI '20· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 60%
Assessment of Sign Language-Based versus Touch-Based Input for Deaf Users Interacting with Intelligent Personal Assistants
CHI '24· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 60%
Tap to Sign: Towards using American Sign Language for text entry on smartphones
MobileHCI '23· Intelligent Voice Assistants (Alexa, Siri, etc.) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)