Show, Not Tell: A Human-AI Collaborative Approach for Designing Sound Awareness Systems

Electrical Muscle Stimulation (EMS)Voice User Interface (VUI) DesignDeaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)Speech-Language Pathologists & AudiologistsAssistive Technology Specialists

Title of the Paper

A Human-AI Collaborative Approach for Designing Sound Awareness Systems

Paper Information

  • Domain: Artificial Intelligence, Human-Computer Interaction, Sound Recognition, Assistive Technology
  • Keywords: AI Collaboration, Sound Classification, Deaf Culture, Sound Perception, Accessible Design, American Sign Language (ASL), Sound Recognition, Assistive Technology, DHH User Experience, Sound Processing

Research Background and Problem

  • Problem or Challenge: Existing sound recognition systems (e.g., SoundWatch) fail to distinguish sounds with similar physical properties (such as the alarm sounds of a microwave and a heart rate monitor) and often misidentify sounds due to a lack of contextual information. Additionally, these systems struggle to adapt to complex sound states (e.g., knocking on a door versus slamming a door).
  • Importance: Providing reliable sound awareness systems is crucial for Deaf and Hard of Hearing (DHH) users, as it helps them better perceive everyday sounds and improves their quality of life.
  • Motivation and Related Work:
    • Current sound classification methods primarily rely on auditory perception, which does not cater to the needs of DHH users.
    • Sound classification approaches include methods based on sound sources, interactivity, signal characteristics, and hybrid features, but they lack classification methods designed from the perspective of DHH users.
    • Existing systems have limitations in accessible design and need to incorporate contextual knowledge from DHH users to improve sound classification accuracy.

Solution

  • Proposed Method or Solution: The authors proposed an innovative "Human-AI Collaborative Sound Awareness System" (HACS). This system combines AI's ability to identify sound characteristics with DHH users' contextual understanding to achieve more accurate sound event recognition.
  • Innovation: HACS compels AI to focus on identifying sound "characteristics" (e.g., "ticking" or "liquid flowing") rather than traditional sound sources or sound events. Users then infer sound events (e.g., a microwave operating in the kitchen) based on these sound characteristics and real-time contextual information.
  • Implementation Steps and Key Technologies:
    • The sound recognition model extracts audio signals from the environment and analyzes their characteristics.
    • The system conveys sound characteristic information to DHH users.
    • Users judge sound events based on sound characteristics and contextual information.
    • An 18-category sound classification table based on ASL was designed, with classifications based on sound characteristics (e.g., "liquid flowing," "machine buzzing"). The classification table was validated through collaboration between AI and DHH users.

Research Outcomes

  • Specific Outcomes:
    1. Proposed an innovative human-AI collaboration framework (HACS) for designing sound awareness systems.
    2. Developed a classification system consisting of 18 sound categories, based on sound characteristics and ASL's visual-spatial expressions.
    3. Validated HACS's potential to assist users in identifying sound events through two preliminary evaluations:
      • Simulated Evaluation: In practical tests, DHH users were able to correctly infer sound events based on independent sound characteristics and contextual information.
      • Algorithm Evaluation: AI models trained on the classification table achieved a classification accuracy of 98.6% on a small dataset.
  • Advantages:
    • Improved recognition accuracy, addressing the challenge of distinguishing similar sounds in traditional sound event recognition systems.
    • Increased autonomy for DHH users in sound processing.
    • High customizability, allowing sound categories to be adjusted based on context or personal habits.
  • Experimental or Evaluation Results:
    • In PE1, test subjects could adapt sound categories to specific sound events in different contexts.
    • In PE2, AI models achieved near-perfect algorithmic recognition performance for the sound categories in the classification table.
  • Limitations and Future Directions:
    • The current classification table may lack comprehensive coverage, such as the exclusion of music or melody-related sounds.
    • Mechanisms for handling overlapping or simultaneous sounds require further research.
    • The system may need to be extended to other sign language systems (e.g., Indian Sign Language, Chinese Sign Language) to meet the needs of a broader user base.
    • Long-term validation of user experience is necessary to ensure the HACS system's effectiveness in diverse environments.

This study represents a significant breakthrough in the field of sound awareness by combining artificial intelligence and user capabilities, pioneering a new approach to accessible design while enhancing the usability and inclusivity of AI-assisted systems.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147942/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642062
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Electrical Muscle Stimulation (EMS), Voice User Interface (VUI) Design, Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)
work
Professions
Speech-Language Pathologists & Audiologists, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
1 related papers