Look Once to Hear: Target Speech Hearing with Noisy Examples
Honorable MentionAuthors
Vibrotactile Feedback & Skin StimulationEye Tracking & Gaze InteractionVoice AccessibilityPhysicians, Nurses & CliniciansSpeech-Language Pathologists & AudiologistsSocial Workers
Document Title
Look Once to Hear: Target Speech Hearing with Noisy Examples
Document Information
- Subject Areas: Human-Computer Interaction, Machine Learning, Smart Wearable Devices
- Keywords: Target Speech Hearing, Noise Processing, Smart Ear-Worn Devices, Real-Time Signal Processing, Deep Learning, Speech Separation, Embedded Systems, Spatial Audio
Research Background and Problem
-
Identified Problems or Challenges:
- In noisy environments, humans can focus on the speech of a target speaker, but existing noise-canceling headphones cannot selectively play specific speech based on the target speaker's voice characteristics.
- Common target speech separation models require clean audio samples of the target speaker as input, which are often unavailable in real-world scenarios.
- User interface design poses challenges, such as how to conveniently obtain registration data for the target speech.
-
Significance:
- Accurately extracting target speech in noisy environments is crucial for daily communication, enhancing auditory experiences, and optimizing hearing aid devices. This technology can not only improve auditory experiences in daily life but also further advance the development of smart wearable devices.
-
Research Motivation:
- To develop a smart wearable device capable of operating effectively in real-world noisy environments, completing target speaker registration in just a few seconds.
- To integrate deep learning and real-time signal processing technologies to optimize device performance, enabling seamless adaptation to various mobile and multipath scenarios.
Solution
-
Proposed Method or Solution:
- A "Target Speech Hearing" system is proposed, which accomplishes the task of target speech separation using two modules:
- Noise Sample Registration Module: Users generate complex signals by briefly gazing at the target speaker (1-4 seconds), which are used to create speaker feature embeddings.
- Real-Time Target Speech Separation Module: An embedded system runs an optimized neural network to separate target speech in real time and remove interfering sounds.
- A "Target Speech Hearing" system is proposed, which accomplishes the task of target speech separation using two modules:
-
Innovations:
- Enables target speaker registration without requiring clean speech samples.
- Optimized neural network architecture reduces runtime latency, allowing real-time operation on embedded devices.
- Designed a training mechanism tailored for real-world scenarios, enabling the model to perform effectively in dynamic environments (e.g., user movement, indoor/outdoor noise).
-
Implementation Steps and Key Technologies:
- Noise Registration Interface:
- Proposed two methods for generating target speaker embeddings: a beamforming-based estimation method and a knowledge distillation estimation method.
- Registration speech is captured as binaural recordings, and speaker embeddings are generated using a speech feature model.
- Real-Time Target Speech Separation System:
- Utilized an improved neural network (e.g., TFGridNet) for target speech separation, combining various techniques to optimize inference time and meet real-time requirements.
- Cached intermediate data to reduce computational redundancy, improving model optimization on embedded platforms.
- Model Training Method:
- Trained the model using synthetic datasets while incorporating BRIR (Binaural Room Impulse Response) characteristics to adapt to real-world scenarios involving multipath and head-related transfer functions.
- Enhanced model robustness by extending training to account for moving sources and real-time errors.
- Noise Registration Interface:
Research Outcomes
-
Specific Results:
- In real-world scenarios, speech signal quality improved by an average of 7.01 dB when using noise registration data, with only a 0.4 dB performance drop compared to using clean data.
- The system processes 8 ms audio chunks in an average of 6.24 ms on embedded devices (e.g., Orange Pi 5B), achieving a total latency of only 18.24 ms.
- User studies demonstrated that the system effectively separates target speakers in indoor and outdoor multipath environments, achieving significant user satisfaction.
-
Advantages Over Existing Solutions:
- Unlike other systems that typically require clean speech registration, this system completes the same task using noisy samples, making it convenient and user-friendly.
- Unlike traditional directional hearing systems, it does not require continuous tracking of the speaker's direction, allowing users to move freely.
- The optimized model runs in real time on embedded systems, achieving low-power, high-efficiency performance.
-
Experimental and Evaluation Results:
- Tests showed that the knowledge distillation network outperformed the beamforming method, delivering better results under noisy registration data conditions.
- In real-world user surveys, over 78% of participants gave high ratings, and the system demonstrated adaptability to rapid user directional changes.
-
Limitations and Future Directions:
- The system currently supports only single-target speech extraction; future work could expand to multi-target speech scenarios.
- The system struggles to separate voices when the target speaker and interferer have highly similar voices, requiring more powerful embedding models and additional registration data.
- The prototype is based on existing headphones and embedded devices; future research could explore more compact and socially acceptable designs (e.g., wireless earbud form factors).
- Further optimization is needed to integrate the device into more advanced active noise-canceling hardware.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How can target speech separation be achieved without clean speech input?Category: Noise-Robust Speech Capture DevicesSimilar questionsarrow_forward
- In noisy environments, how can target speaker enrollment be completed quickly for speech separation?Category: Noise-Robust Speech Capture DevicesSimilar questionsarrow_forward
- How can real-time speech separation systems be optimized for mobile scenarios and multipath environments?Category: Noise-Robust Speech Capture DevicesSimilar questionsarrow_forward
lightbulb
Practical Problems
1- In noisy environments, users cannot clearly hear speech from a designated speaker.Category: Noise-Robust Speech Capture DevicesSimilar questionsarrow_forward
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642057
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
Honorable Mention
group
Authors
5 authors
sell
Subtopics
Vibrotactile Feedback & Skin Stimulation, Eye Tracking & Gaze Interaction, Voice Accessibility
work
Professions
Physicians, Nurses & Clinicians, Speech-Language Pathologists & Audiologists, Social Workers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers