SceneScout: Towards AI-Driven Access to Street Level Imagery for Blind Users

Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)Privacy by Design & User ControlGenerative AI (Text, Image, Music, Video)Human-LLM CollaborationPhysicians, Nurses & CliniciansSpeech-Language Pathologists & AudiologistsAssistive Technology Specialists

Paper Title

SceneScout: Towards AI-Driven Access to Street Level Imagery for Blind Users

Publication Info

  • Topic area: AI-driven accessibility solutions for blind and low-vision (BLV) users in navigation.
  • Keywords: Blind users, low-vision, street level imagery, multimodal large language models, navigation assistance, accessibility, route preview, virtual exploration, spatial reasoning, user study.

Background and Problem

  • Problem / challenge: Existing navigation tools for BLV users focus on in-situ guidance but lack detailed pre-travel information about environmental accessibility features. Street level imagery, rich in visual context, remains inaccessible to BLV users due to its visual and spatial complexity.
  • Significance: Providing BLV users with access to street level imagery could enhance their confidence and independence in navigating unfamiliar environments by offering detailed spatial and environmental insights.
  • Motivation and related work: Previous research has explored tactile maps, audio-haptic feedback, and virtual navigation systems but often overlooks environmental details like curb cuts or tactile paving. While prior work has used street level imagery for accessibility assessments, it has not been made directly usable by BLV users. This paper addresses this gap by enabling BLV users to interact with street level imagery independently.

Solution

  • Proposed approach: SceneScout, a multimodal large language model (MLLM)-driven prototype, enables BLV users to access and interpret street level imagery through two interaction modes: Route Preview and Virtual Exploration.
  • Novelty:
    1. Development of a system that integrates MLLMs with street level imagery to generate spatially grounded, personalized descriptions.
    2. Introduction of two interaction modes tailored to BLV users: Route Preview for structured route narratives and Virtual Exploration for open-ended neighborhood exploration.
    3. Evaluation of MLLM-generated descriptions for accuracy, relevance, and temporal consistency in a user study with BLV participants.
  • Procedure and key techniques:
    • Route Preview: Fuses successive panoramas into structured narratives with high-level overviews and mobility-critical details.
    • Virtual Exploration: Allows users to specify exploration goals and navigate street level imagery interactively.
    • Computational pipeline: Segments panoramas into directionally meaningful views, integrates map metadata, and generates orientation-aware descriptions tailored to user intent.

Results

  • Concrete findings:
    • 72% of descriptions were accurate, 95% were likely to remain consistent over time, and 96% of Virtual Exploration descriptions aligned with user intent.
    • Errors included plausible but unverifiable details (40%), factual inaccuracies (19%), spatial errors (16%), and hallucinations (16%).
  • Advantage over baselines:
    • SceneScout provided richer environmental details compared to existing navigation tools, enabling BLV users to uncover visual information otherwise inaccessible.
    • Participants valued the ability to specify keywords and receive personalized descriptions.
  • Experiments / evaluation:
    • Mixed-methods user study with 10 BLV participants.
    • Scenarios included familiar and unfamiliar locations for both interaction modes.
    • Metrics assessed: relevance, utility, trust, confidence, temporal consistency, and redundancy.
  • Limitations and future work:
    • Spatial imprecision and occasional assumptions about user abilities reduced trust.
    • Limited diversity in participant demographics and lack of real-world navigation testing.
    • Future work should focus on refining spatial reasoning, enabling backtracking, and integrating real-time navigation assistance.

Summary

SceneScout is a prototype system that leverages multimodal large language models to make street level imagery accessible to BLV users. It supports two interaction modes—Route Preview and Virtual Exploration—enabling users to access detailed spatial and environmental information. A user study demonstrated that SceneScout effectively surfaces relevant details, fostering confidence and independence in navigation. However, challenges such as spatial imprecision, temporal inconsistencies, and assumptions about user capabilities highlight areas for improvement. This work marks an initial step toward AI-driven accessibility solutions for BLV users, with potential applications in pre-travel planning and real-time navigation.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223523/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790449
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille), Privacy by Design & User Control, Generative AI (Text, Image, Music, Video), Human-LLM Collaboration
work
Professions
Physicians, Nurses & Clinicians, Speech-Language Pathologists & Audiologists, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
0 related papers