(Computer) Vision in Action: Comparing Remote Sighted Assistance and a Multimodal Voice Agent in Inspection Sequences

Intelligent Voice Assistants (Alexa, Siri, etc.)Voice AccessibilitySpeech-Language Pathologists & AudiologistsHCI Researchers

Paper Title

(Computer) Vision in Action: Comparing Remote Sighted Assistance and a Multimodal Voice Agent in Inspection Sequences

Publication Info

  • Topic area: Comparative analysis of human and AI-powered assistance in assistive tasks for blind individuals.
  • Keywords: remote sighted assistance, multimodal voice agents, proactivity, human-AI collaboration, assistive technology, ethnomethodology, conversation analysis, blind and low vision, collaborative practices, turn-taking.

Background and Problem

  • Problem / challenge: Current multimodal voice agents lack the ability to initiate and modify actions based on real-time environmental cues, limiting their effectiveness in collaborative tasks compared to human assistants.
  • Significance: Addressing this gap is crucial for improving assistive technologies for blind and low-vision (BLV) individuals, who rely on such tools for everyday tasks like cleanliness inspection.
  • Motivation and related work: While prior studies have explored the capabilities of multimodal voice agents and remote sighted assistance, they have not provided fine-grained analyses of the collaborative practices that underpin successful human assistance. This paper aims to fill that gap by comparing interactions involving a human assistant and a voice agent.

Solution

  • Proposed approach: A comparative ethnomethodological conversation analysis of interactions between a blind participant and two types of assistants: a human remote sighted assistant (RSA) and a multimodal voice agent (ChatGPT multimodal “voice mode”).
  • Novelty:
    1. Detailed identification of proactive practices in human assistance that are absent in AI-powered assistance.
    2. Empirical specification of the conditions under which assistive collaboration succeeds or fails.
    3. Ethical discussion of the implications of enabling voice agents to initiate environmentally occasioned actions.
  • Procedure and key techniques:
    • Data collection involved video-recorded interactions of a blind participant performing a stain-finding task with both a human RSA and a multimodal voice agent.
    • Analysis used ethnomethodological conversation analysis (EMCA) to examine turn-taking, mutual adjustments, and task coordination.
    • Transcripts were created using Jeffersonian and Mondadian conventions to capture speech and embodied actions.

Results

  • Concrete findings:
    • The human RSA initiated actions based on visual cues, modified actions in real-time, and distributed perceptual work with the participant, leading to task success.
    • The voice agent failed to initiate or adapt actions based on environmental cues, resulting in a rigid, step-by-step interaction structure that did not locate the stain.
  • Advantage over baselines:
    • Human RSA demonstrated superior proactivity and adaptability, enabling effective collaboration and task completion.
    • Voice agents lacked the ability to respond to or initiate actions based on real-time visual data, highlighting a significant limitation.
  • Experiments / evaluation:
    • Two fragments from larger corpora were analyzed: one involving a human RSA using the Be My Eyes app and another involving ChatGPT multimodal “voice mode”.
    • Metrics included the ability to locate the stain, the nature of turn-taking, and the distribution of perceptual work.
  • Limitations and future work:
    • The study's findings are based on a small sample size and specific configurations of the voice agent, limiting generalizability.
    • Future research should explore a broader range of tasks, settings, and assistive technologies, as well as the ethical implications of proactive AI behaviors.

Summary

This study compared the collaborative practices of a human remote sighted assistant and a multimodal voice agent in assisting a blind participant with a stain-finding task. The human assistant's success relied on proactive behaviors, including initiating actions based on visual cues, real-time adjustments, and distributing perceptual work. In contrast, the voice agent's rigid, step-by-step interaction failed to achieve the task. The findings highlight the importance of proactivity and environmentally responsive actions for effective assistive technologies. However, the study also raises ethical concerns about enabling AI to autonomously initiate actions, as this could impose value judgments on users. Future work should address these ethical and technical challenges to improve the design of assistive voice agents.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222431/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791708
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Intelligent Voice Assistants (Alexa, Siri, etc.), Voice Accessibility
work
Professions
Speech-Language Pathologists & Audiologists, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
4 related papers