Beyond Descriptions: A Generative Scene2Audio Framework for Blind and Low-Vision Users to Experience Vista Landscapes
Authors
Paper Title
Beyond Descriptions: A Generative Scene2Audio Framework for Blind and Low-Vision Users to Experience Vista Landscapes
Publication Info
- Topic area: Assistive technologies for Blind and Low-Vision (BLV) users, focusing on auditory scene representation.
- Keywords: Scene sonification, generative audio, assistive technology, blind and low-vision, auditory scene analysis, psychoacoustics, user experience, cognitive load, accessibility, Vista spaces.
Background and Problem
- Problem / challenge: Current assistive technologies for BLV users rely heavily on verbal descriptions, which are transactional, impose cognitive load, and fail to evoke the aesthetic pleasure of visual experiences. Existing generative models struggle to effectively capture and sonify complex scenes with multiple objects.
- Significance: Enhancing the experience of distant scenic views (Vista spaces) for BLV users is crucial for leisure, emotional well-being, and cognitive engagement. Addressing the gap between visual and auditory scene perception can improve accessibility and quality of life.
- Motivation and related work: Prior work has shown that non-verbal audio enhances immersion and comprehension in storytelling and media. However, existing assistive tools for BLV users are limited to image-to-text-to-speech frameworks, which lack engagement and scalability. Generative models for audio synthesis often fail in scene-level composition, especially for complex environments.
Solution
- Proposed approach: Scene2Audio framework, which generates non-verbal soundscapes for Vista spaces by combining AI-driven object sonification with psychoacoustic principles and auditory scene analysis.
- Novelty:
- Introduction of a two-step process: Salient Objects Identification and Audio Scene Composition.
- Integration of psychoacoustics and Foley sound synthesis for realistic and immersive auditory scenes.
- Validation through lab and in-the-wild studies with BLV participants, demonstrating ecological validity.
- Procedure and key techniques:
- Salient Objects Identification: Uses GPT-4 to identify sound-making objects in a scene and generate noun-verb action phrases (e.g., "leaves rustling").
- Audio Scene Composition: Differentiates discrete and continuous sounds, minimizes repetitive discrete events, and layers sounds with weighted mixing (foreground: 0.8, background: 0.2).
- Evaluation through controlled lab studies and real-world app deployment with BLV users.
Results
- Concrete findings:
- Scene2Audio achieved an average image recognition accuracy of 62.4%, outperforming baselines like Im2wav (17.3%) and Im2text2audio (29.6%).
- Overlay audio (speech + non-verbal sounds) was the most preferred mode in lab studies, scoring highest in comprehension (6.11 ± 0.53), immersion (5.03 ± 1.00), and cognitive load (2.04 ± 1.01).
- In real-world app usage, the detail mode (Overlay-Concat) was preferred for its clarity (92%) and enjoyability (88%).
- Advantage over baselines:
- Scene2Audio captured complex scenes more effectively than Im2wav and Im2text2audio, particularly in nature environments.
- Higher confidence (5.48 ± 1.15) and pleasantness (5.14 ± 1.24) ratings compared to baselines.
- Experiments / evaluation:
- Lab study with 11 BLV participants evaluated four audio feedback strategies (Speech-only, Audio-only, Overlay, Overlay-Concat) across eight Vista scenes.
- In-the-wild study with 7 BLV participants using a mobile app over a week, collecting feedback on 77 images.
- Listening test with 21 sighted participants compared Scene2Audio with baseline methods.
- Limitations and future work:
- Small sample size of BLV participants limits generalizability.
- Challenges in controlling discrete sound events and addressing hallucinations in generative models.
- Practical deployment issues include latency, lack of spatialization, and need for context-aware modes.
- Future work includes integrating spatial audio, reducing latency, and improving generative model control.
Summary
The Scene2Audio framework bridges the gap between visual and auditory scene perception for BLV users by generating non-verbal soundscapes informed by psychoacoustics and auditory scene analysis. It significantly outperforms baseline methods in comprehensibility and user preference, particularly for nature scenes. Lab and in-the-wild studies demonstrate its potential to enhance leisure and cognitive engagement in Vista spaces. While challenges remain in scalability, latency, and real-world deployment, this work lays the foundation for accessible and immersive auditory experiences in assistive technologies.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)