Weaving Sound Information to Support Real-Time Sensemaking of Auditory Environments: Co-Designing with a DHH User

Voice AccessibilityDeaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)Universal & Inclusive DesignSpeech-Language Pathologists & AudiologistsAssistive Technology Specialists

Research Background and Issues

  • What problems or challenges did the authors identify?
    Current artificial intelligence (AI) sound recognition systems assist Deaf and Hard of Hearing (DHH) individuals in accessing sound information, such as discrete sound sources and speech transcription. However, these systems are often pre-configured and cannot be customized to accommodate the dynamic real-time contexts, goals, and information needs of DHH users. This static design struggles to meet users' dynamic needs in complex auditory environments and support semantic understanding.

  • Why is this issue important?
    DHH users rely on sound awareness technologies to better understand their surroundings, thereby supporting daily social interactions and task execution. Dynamically adjusting the display of sound information to meet real-time needs not only enhances user experience but also improves DHH users' environmental understanding, promoting the adoption of accessible technologies.

  • Research Motivation and Related Work
    Previous work has primarily focused on technologies such as sound classification, auditory scene understanding, and automatic speech recognition. However, the information generated is often static and pre-configured, lacking semantic relevance to user goals. The authors aim to overcome the limitations of traditional AI sound recognition systems by designing an "intention-driven" system that weaves sound information and dynamically adjusts it in real time based on users' evolving needs.

Solution

  • What methods or solutions did the authors propose?
    The authors proposed a prototype system called "SoundWeaver," which dynamically weaves sound outputs from different AI models based on user intentions and presents integrated information through a head-mounted display (HMD). The system includes three modes: Awareness (environmental perception), Action (task monitoring), and Social (social interaction), each catering to different information needs of DHH users.

  • What are the innovative aspects of this solution?
    The primary distinction between SoundWeaver and traditional sound recognition systems lies in its "intention-driven" design. Unlike static visual outputs, this system dynamically adjusts its behavior based on the user's current goals, supporting real-time semantic perception and switching. Furthermore, it emphasizes user-specific environments and cultural contexts through multiple co-design iterations, avoiding overly intrusive designs.

  • What are the implementation steps and key technologies used?

    1. Multi-stage Co-design: The authors conducted a three-stage study involving in-depth co-design with a DHH user ("Declan") to iteratively refine the system.
    2. Multimodal AI Support: The system integrates technologies such as sound classification, speech recognition, and acoustic scene understanding (e.g., the Audio Flamingo model) to dynamically generate audio and semantic information.
    3. Head-mounted Display Implementation: SoundWeaver operates on the Apple Vision Pro device, enabling users to switch modes and adjust information displays through simple interactions.
    4. Mode Design:
      • Awareness mode provides an overall perception of environmental sounds, such as visualizing volume changes through waveforms.
      • Action mode supports task-related sound monitoring, allowing users to track specific sounds, such as a microwave's completion alert.
      • Social mode assists users in social interactions through speech transcription and social indicators (e.g., "sound bubbles" displays).

Research Outcomes

  • What specific outcomes were achieved?
    SoundWeaver streamlines different types of sound information through its three modes, enabling users to switch between them in real time based on their needs. Additionally, the system includes an overlay feature for salient sounds (e.g., emergency alarms and name calls), ensuring timely alerts regardless of the current mode.

  • What advantages does it have compared to existing solutions?
    Compared to the static design of existing systems, SoundWeaver offers dynamic information scheduling based on individual intentions and significantly optimizes sound visualization to avoid interference or information overload. The system also incorporates users' cultural contexts and personalized needs, achieving a more tailored user experience through co-design.

  • What were the experimental or evaluation results?
    Deployments and evaluations in two real-world environments (a home and a gaming store) demonstrated that SoundWeaver effectively supports users in task execution and social interactions. For example, users could easily switch information displays based on real-time needs through dynamic mode switching. The system also facilitated new social interaction dynamics, such as enhancing collaboration with friends.

  • Limitations and Future Directions

    1. Limited User Diversity: The current study is based on the deep involvement of a single participant, Declan. Future research should expand to a broader DHH population to validate the design's applicability.
    2. Hardware Comfort: The current device (e.g., Apple Vision Pro) is relatively bulky, causing fatigue during prolonged use. Future work should explore more user-friendly hardware forms.
    3. Implicit Feedback Design: Currently, mode switching requires manual operation. Future research could explore more intelligent intention inference technologies to reduce interaction complexity.
    4. Potential Conflicts with Users' Social Dynamics: The system design should carefully address users' existing social relationships to avoid potential intrusiveness.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188590/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714268
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Voice Accessibility, Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration), Universal & Inclusive Design
work
Professions
Speech-Language Pathologists & Audiologists, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
9 related papers