Look Once to Hear: Target Speech Hearing with Noisy Examples

Honorable Mention
Vibrotactile Feedback & Skin StimulationEye Tracking & Gaze InteractionVoice AccessibilityPhysicians, Nurses & CliniciansSpeech-Language Pathologists & AudiologistsSocial Workers

Document Title

Look Once to Hear: Target Speech Hearing with Noisy Examples

Document Information

  • Subject Areas: Human-Computer Interaction, Machine Learning, Smart Wearable Devices
  • Keywords: Target Speech Hearing, Noise Processing, Smart Ear-Worn Devices, Real-Time Signal Processing, Deep Learning, Speech Separation, Embedded Systems, Spatial Audio

Research Background and Problem

  • Identified Problems or Challenges:

    • In noisy environments, humans can focus on the speech of a target speaker, but existing noise-canceling headphones cannot selectively play specific speech based on the target speaker's voice characteristics.
    • Common target speech separation models require clean audio samples of the target speaker as input, which are often unavailable in real-world scenarios.
    • User interface design poses challenges, such as how to conveniently obtain registration data for the target speech.
  • Significance:

    • Accurately extracting target speech in noisy environments is crucial for daily communication, enhancing auditory experiences, and optimizing hearing aid devices. This technology can not only improve auditory experiences in daily life but also further advance the development of smart wearable devices.
  • Research Motivation:

    • To develop a smart wearable device capable of operating effectively in real-world noisy environments, completing target speaker registration in just a few seconds.
    • To integrate deep learning and real-time signal processing technologies to optimize device performance, enabling seamless adaptation to various mobile and multipath scenarios.

Solution

  • Proposed Method or Solution:

    • A "Target Speech Hearing" system is proposed, which accomplishes the task of target speech separation using two modules:
      1. Noise Sample Registration Module: Users generate complex signals by briefly gazing at the target speaker (1-4 seconds), which are used to create speaker feature embeddings.
      2. Real-Time Target Speech Separation Module: An embedded system runs an optimized neural network to separate target speech in real time and remove interfering sounds.
  • Innovations:

    • Enables target speaker registration without requiring clean speech samples.
    • Optimized neural network architecture reduces runtime latency, allowing real-time operation on embedded devices.
    • Designed a training mechanism tailored for real-world scenarios, enabling the model to perform effectively in dynamic environments (e.g., user movement, indoor/outdoor noise).
  • Implementation Steps and Key Technologies:

    1. Noise Registration Interface:
      • Proposed two methods for generating target speaker embeddings: a beamforming-based estimation method and a knowledge distillation estimation method.
      • Registration speech is captured as binaural recordings, and speaker embeddings are generated using a speech feature model.
    2. Real-Time Target Speech Separation System:
      • Utilized an improved neural network (e.g., TFGridNet) for target speech separation, combining various techniques to optimize inference time and meet real-time requirements.
      • Cached intermediate data to reduce computational redundancy, improving model optimization on embedded platforms.
    3. Model Training Method:
      • Trained the model using synthetic datasets while incorporating BRIR (Binaural Room Impulse Response) characteristics to adapt to real-world scenarios involving multipath and head-related transfer functions.
      • Enhanced model robustness by extending training to account for moving sources and real-time errors.

Research Outcomes

  • Specific Results:

    • In real-world scenarios, speech signal quality improved by an average of 7.01 dB when using noise registration data, with only a 0.4 dB performance drop compared to using clean data.
    • The system processes 8 ms audio chunks in an average of 6.24 ms on embedded devices (e.g., Orange Pi 5B), achieving a total latency of only 18.24 ms.
    • User studies demonstrated that the system effectively separates target speakers in indoor and outdoor multipath environments, achieving significant user satisfaction.
  • Advantages Over Existing Solutions:

    • Unlike other systems that typically require clean speech registration, this system completes the same task using noisy samples, making it convenient and user-friendly.
    • Unlike traditional directional hearing systems, it does not require continuous tracking of the speaker's direction, allowing users to move freely.
    • The optimized model runs in real time on embedded systems, achieving low-power, high-efficiency performance.
  • Experimental and Evaluation Results:

    • Tests showed that the knowledge distillation network outperformed the beamforming method, delivering better results under noisy registration data conditions.
    • In real-world user surveys, over 78% of participants gave high ratings, and the system demonstrated adaptability to rapid user directional changes.
  • Limitations and Future Directions:

    • The system currently supports only single-target speech extraction; future work could expand to multi-target speech scenarios.
    • The system struggles to separate voices when the target speaker and interferer have highly similar voices, requiring more powerful embedding models and additional registration data.
    • The prototype is based on existing headphones and embedded devices; future research could explore more compact and socially acceptable designs (e.g., wireless earbud form factors).
    • Further optimization is needed to integrate the device into more advanced active noise-canceling hardware.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147319/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642057
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
Honorable Mention
group
Authors
5 authors
sell
Subtopics
Vibrotactile Feedback & Skin Stimulation, Eye Tracking & Gaze Interaction, Voice Accessibility
work
Professions
Physicians, Nurses & Clinicians, Speech-Language Pathologists & Audiologists, Social Workers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers