Semantic Hearing: Programming Acoustic Scenes with Binaural Hearables

Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)Biosensors & Physiological MonitoringPhysicians, Nurses & CliniciansAthletes & Fitness EnthusiastsAssistive Technology Specialists

Title of the Paper

Semantic Hearing: Programming Acoustic Scenes with Binaural Hearables

Paper Information

  • Subject Area: Audio processing and human-computer interaction, focusing on semantic hearing and real-time sound extraction with binaural devices
  • Keywords: Binaural target sound extraction, embedded computing, noise cancellation, real-time semantic hearing, human-computer interaction, stochastic neural networks, spatial audio, deep learning

Research Background and Problem

  • Problem or Challenge:
    • In real-world environments, users may want to focus on specific audio targets in real time (e.g., bird songs, ambulance sirens) while filtering out other distractions (e.g., traffic noise). Traditional noise cancellation devices can only process environmental noise uniformly and cannot flexibly select which sounds to allow or block. Achieving this functionality faces challenges such as real-time processing based on binaural input, efficiently preserving spatial audio effects, and the model's limited generalization ability in complex real-world scenarios.
  • Significance:
    • Semantic hearing enables programmable control over acoustic scenes for wearable devices. This not only enhances users' audio experience but also has applications in health, entertainment, and safety, such as clearly focusing on a baby's cry or a police siren.
  • Motivation and Related Work:
    • Noise cancellation technology is mature but only filters all environmental sounds or performs simple transparency processing. In recent years, target sound extraction techniques have made progress but are mostly focused on offline speech separation or single-channel models, with limited research addressing real-time scenarios for binaural devices.

Solution

  • Research Methods and Solution:
    • Proposed the first neural network architecture for target sound extraction based on binaural input, capable of effectively preserving spatial audio cues during real-time operation.
    • Designed an improved Transformer network, combined with dynamic programming methods to optimize processing efficiency, enabling real-time operation on low-end devices (e.g., smartphones).
    • Developed a data generation and model training method that synthesizes data covering various room reverberation responses and head-related transfer functions (HRTFs), significantly improving the model's generalization ability.
  • Innovations:
    • The first binaural framework to jointly process left and right ear channels instead of handling them separately, reducing computational cost by 50%.
    • Training methods using synthetic data that incorporate real-world factors (e.g., reverberation, sound reflections), avoiding reliance on hardware-captured data and improving performance in new environments and with new users.
  • Key Implementation Steps and Techniques:
    1. Designed a binaural sound extraction network architecture suitable for real-time requirements, including encoder and decoder modules.
    2. Implemented dynamic streaming inference while maintaining logical causality.
    3. Dynamically synthesized binaural mixed data with head-related transfer functions using the Scaper tool.
    4. Trained the model using various audio datasets and synthesized sound scenes to enhance its ability to recognize diverse environments and target audio effects.

Research Outcomes

  • Specific Results:
    • The model achieved an average signal quality improvement [SNR] of 7.17 dB for extracting 20 types of target sounds in real-world environments.
    • The system demonstrated real-time capabilities, with a processing time of only 6.56ms for 10ms audio chunks on an iPhone 11, meeting real-time playback requirements.
    • Users were able to adjust the device to block interfering sounds and significantly enhance the auditory experience of target sounds.
  • Advantages:
    • Compared to traditional noise cancellation devices, this solution allows for semantic control to selectively listen to target sounds while preserving spatial awareness and sound source direction.
    • Compared to existing sound extraction methods, the model is more efficient and has better generalization capabilities.
  • Experimental or Evaluation Results:
    • User experiments conducted in five real-world complex environments demonstrated that the model accurately preserved the perceived sound source direction, with a 50th percentile angular error of 22.5° and a 90th percentile error of 45°.
    • Users preferred using a voice user interface to select target sounds, demonstrating the system's ease of use and practicality.
  • Limitations and Future Directions:
    • Data imbalance: Certain target sounds (e.g., car horns) have fewer training samples, affecting performance.
    • Difficulty in separating sounds with similar characteristics (e.g., music and speech), requiring improvements in network architecture or increased data diversity.
    • Current hardware integration requires more compact devices, such as headphones optimized for both noise cancellation and playback functions.
    • In the long term, integrating custom silicon chips to reduce device power consumption and latency is a potential direction for commercialization.

Conclusion and Contributions

This study is the first to propose the concept of real-time semantic hearing and implement a prototype system, pioneering the solution to selective sound extraction in real-world scenarios for binaural devices. By releasing public datasets and code, the authors aim to promote further research and applications in this field.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/uist/126683/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3586183.3606779
At a Glance

Paper Snapshot

fact_check
dataset
Source
UIST
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration), Biosensors & Physiological Monitoring
work
Professions
Physicians, Nurses & Clinicians, Athletes & Fitness Enthusiasts, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
0 related papers