Beyond Descriptions: A Generative Scene2Audio Framework for Blind and Low-Vision Users to Experience Vista Landscapes

Audio Accessibility (Captions, Sign Language, Vibration)Emotion-Sensing WearablesVisual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)Speech-Language Pathologists & AudiologistsElderly Care WorkersFamily Caregivers

Paper Title

Beyond Descriptions: A Generative Scene2Audio Framework for Blind and Low-Vision Users to Experience Vista Landscapes

Publication Info

  • Topic area: Assistive technologies for Blind and Low-Vision (BLV) users, focusing on auditory scene representation.
  • Keywords: Scene sonification, generative audio, assistive technology, blind and low-vision, auditory scene analysis, psychoacoustics, user experience, cognitive load, accessibility, Vista spaces.

Background and Problem

  • Problem / challenge: Current assistive technologies for BLV users rely heavily on verbal descriptions, which are transactional, impose cognitive load, and fail to evoke the aesthetic pleasure of visual experiences. Existing generative models struggle to effectively capture and sonify complex scenes with multiple objects.
  • Significance: Enhancing the experience of distant scenic views (Vista spaces) for BLV users is crucial for leisure, emotional well-being, and cognitive engagement. Addressing the gap between visual and auditory scene perception can improve accessibility and quality of life.
  • Motivation and related work: Prior work has shown that non-verbal audio enhances immersion and comprehension in storytelling and media. However, existing assistive tools for BLV users are limited to image-to-text-to-speech frameworks, which lack engagement and scalability. Generative models for audio synthesis often fail in scene-level composition, especially for complex environments.

Solution

  • Proposed approach: Scene2Audio framework, which generates non-verbal soundscapes for Vista spaces by combining AI-driven object sonification with psychoacoustic principles and auditory scene analysis.
  • Novelty:
    1. Introduction of a two-step process: Salient Objects Identification and Audio Scene Composition.
    2. Integration of psychoacoustics and Foley sound synthesis for realistic and immersive auditory scenes.
    3. Validation through lab and in-the-wild studies with BLV participants, demonstrating ecological validity.
  • Procedure and key techniques:
    1. Salient Objects Identification: Uses GPT-4 to identify sound-making objects in a scene and generate noun-verb action phrases (e.g., "leaves rustling").
    2. Audio Scene Composition: Differentiates discrete and continuous sounds, minimizes repetitive discrete events, and layers sounds with weighted mixing (foreground: 0.8, background: 0.2).
    3. Evaluation through controlled lab studies and real-world app deployment with BLV users.

Results

  • Concrete findings:
    • Scene2Audio achieved an average image recognition accuracy of 62.4%, outperforming baselines like Im2wav (17.3%) and Im2text2audio (29.6%).
    • Overlay audio (speech + non-verbal sounds) was the most preferred mode in lab studies, scoring highest in comprehension (6.11 ± 0.53), immersion (5.03 ± 1.00), and cognitive load (2.04 ± 1.01).
    • In real-world app usage, the detail mode (Overlay-Concat) was preferred for its clarity (92%) and enjoyability (88%).
  • Advantage over baselines:
    • Scene2Audio captured complex scenes more effectively than Im2wav and Im2text2audio, particularly in nature environments.
    • Higher confidence (5.48 ± 1.15) and pleasantness (5.14 ± 1.24) ratings compared to baselines.
  • Experiments / evaluation:
    • Lab study with 11 BLV participants evaluated four audio feedback strategies (Speech-only, Audio-only, Overlay, Overlay-Concat) across eight Vista scenes.
    • In-the-wild study with 7 BLV participants using a mobile app over a week, collecting feedback on 77 images.
    • Listening test with 21 sighted participants compared Scene2Audio with baseline methods.
  • Limitations and future work:
    • Small sample size of BLV participants limits generalizability.
    • Challenges in controlling discrete sound events and addressing hallucinations in generative models.
    • Practical deployment issues include latency, lack of spatialization, and need for context-aware modes.
    • Future work includes integrating spatial audio, reducing latency, and improving generative model control.

Summary

The Scene2Audio framework bridges the gap between visual and auditory scene perception for BLV users by generating non-verbal soundscapes informed by psychoacoustics and auditory scene analysis. It significantly outperforms baseline methods in comprehensibility and user preference, particularly for nature scenes. Lab and in-the-wild studies demonstrate its potential to enhance leisure and cognitive engagement in Vista spaces. While challenges remain in scalability, latency, and real-world deployment, this work lays the foundation for accessible and immersive auditory experiences in assistive technologies.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223150/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791655
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Audio Accessibility (Captions, Sign Language, Vibration), Emotion-Sensing Wearables, Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
work
Professions
Speech-Language Pathologists & Audiologists, Elderly Care Workers, Family Caregivers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers