SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures
Authors
Paper Title
SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures
Publication Info
- Topic area: Directional speech extraction on smartphones using bio-inspired acoustic microstructures.
- Keywords: directional speech extraction, acoustic microstructure, smartphones, real-time processing, neural networks, spatial audio, beamforming, microphone arrays, user interface, audio enhancement.
Background and Problem
- Problem / challenge: Existing systems for directional speech extraction rely on microphone arrays with multiple channels, which are not compatible with most smartphones due to hardware constraints. Smartphones typically have only two microphones, limiting their ability to capture spatial audio cues. Current directional microphones are fixed-direction and lack multi-directional separation capabilities.
- Significance: Enabling smartphones to perform directional speech extraction would improve audio quality in noisy environments, benefiting applications like meeting transcription, remote conferencing, and voice assistants.
- Motivation and related work: Prior work on spatial audio sensing has focused on microphone arrays or custom hardware, which are incompatible with smartphones. Systems like Owlet and EarCase introduced acoustic microstructures for direction-of-arrival estimation but did not address directional speech extraction. Neural beamformers and statistical algorithms have limitations in handling complex spatial cues. This paper builds on these works by introducing a smartphone-compatible solution that uses a bio-inspired acoustic microstructure.
Solution
- Proposed approach: SonicSieve, a system combining a bio-inspired acoustic microstructure and a real-time neural network to enable directional speech extraction on smartphones.
- Novelty:
- Development of a compact acoustic microstructure optimized for speech frequencies, providing enhanced spatial diversity.
- Integration of the microstructure with standard wired earphones for smartphone compatibility.
- Real-time neural network for directional speech extraction, supporting multi-speaker scenarios.
- Creation of a real-world dataset for training and evaluation across diverse environments.
- Procedure and key techniques:
- Design and fabrication of a 20 mm diameter microstructure with six strategically placed holes to maximize spatial diversity.
- Integration with smartphones using an in-line microphone and alignment with the built-in microphone.
- Development of a neural network leveraging interaural level and phase differences for real-time processing.
- Collection of a real-world dataset in multiple environments to train and evaluate the system.
Results
- Concrete findings:
- Achieved a 5.0 dB SI-SDR improvement for single-sector directional speech extraction, outperforming a baseline without the microstructure (2.3 dB).
- Demonstrated performance comparable to or better than 5-microphone arrays, achieving 4.4 dB SI-SDR improvement with only two microphones.
- Real-time processing capability with latencies of 4.5–7.2 ms across various smartphones.
- User study results showed higher mean opinion scores for SonicSieve compared to a 2-microphone array baseline.
- Advantage over baselines:
- Outperformed traditional 2-microphone systems and matched or exceeded 5-microphone array performance.
- Demonstrated robustness across diverse environments and smartphone geometries.
- Experiments / evaluation:
- Evaluated in five real-world environments with 45,000 audio mixtures.
- Conducted leave-one-room-out validation and user studies with 44 participants.
- Compared against conventional beamforming and microphone array systems.
- Limitations and future work:
- Current system requires manual sector updates for moving speakers.
- Limited to predefined directional sectors; does not separate multiple speakers within the same sector.
- Future work includes dynamic direction-of-arrival estimation, finer spatial separation, and integration into device-native designs.
Summary
SonicSieve introduces a novel approach to directional speech extraction on smartphones by combining a bio-inspired acoustic microstructure with a real-time neural network. The system achieves significant improvements in speech quality (up to 5.0 dB SI-SDR) and matches or exceeds the performance of 5-microphone arrays while using only two microphones. Evaluations across diverse environments and user studies confirm its robustness and usability. The system is compatible with standard smartphones and supports applications like meeting transcription and voice assistants. Future work will address dynamic speaker tracking and finer spatial resolution.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)