NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech Interaction
Paper Title
NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech Interaction
Publication Info
- Topic area: Silent and whispered speech interfaces for AI voice interaction
- Keywords: silent speech, whispered speech, smart glasses, vibration sensor, microphone, noise robustness, always-available interface, wearable technology, speech recognition, multimodal fusion
Background and Problem
- Problem / challenge: Existing silent and whispered speech systems fail to balance wearability, silence, noise robustness, and vocabulary size, limiting their practicality for continuous AI interaction.
- Significance: A robust, unobtrusive, and noise-tolerant speech interface would enable discreet and always-available AI conversations, enhancing usability in public and noisy environments.
- Motivation and related work: Prior systems like neuromuscular sensors, lip-reading, and bone conduction devices have limitations in vocabulary size, noise robustness, or wearability. Whispered speech systems offer potential but are highly susceptible to noise. This paper addresses these gaps by proposing a novel multimodal approach.
Solution
- Proposed approach: NasoVoce, a nose-mounted interface integrating a MEMS microphone and a MEMS vibration sensor to capture both air- and bone-conducted speech signals.
- Novelty:
- Integration of a microphone and vibration sensor on the nasal bridge for capturing normal and whispered speech with noise robustness.
- Development of a dual-input deep learning model (D-DCCRN) for fusing microphone and vibration sensor signals to enhance speech quality and recognition accuracy.
- Creation of a dataset and training methodology for whispered and normal speech in noisy environments.
- Demonstration of the system’s practicality through objective metrics, subjective ratings, and real-world evaluations.
- Procedure and key techniques:
- Hardware: A MEMS microphone and vibration sensor mounted on the nose pads of smart glasses, capturing synchronized audio and vibration signals.
- Software: A D-DCCRN model processes the real and imaginary components of both signals for noise-robust speech enhancement.
- Training: Dataset of 104 hours of paired microphone and vibration sensor recordings with added noise; loss functions include audio enhancement and knowledge distillation.
- Evaluation: ASR accuracy, objective quality metrics (PESQ, STOI), subjective ratings (MUSHRA), and real-world tests.
Results
- Concrete findings:
- D-DCCRN has 438.25M parameters and processes speech in 136.9 ms, requiring only 31.8% of OpenAI Whisper’s processing time.
- Enhanced speech outperformed standalone microphone and vibration sensor inputs in ASR accuracy, PESQ, and STOI metrics under most noise conditions.
- MUSHRA ratings showed enhanced speech consistently scored higher than microphone or vibration-only inputs under various noise levels.
- Advantage over baselines:
- Enhanced speech recognition accuracy was superior to microphone-only and vibration-only inputs, particularly in noisy environments.
- Outperformed Apple AirPods Pro 2’s voice isolation feature in capturing whispered speech.
- Experiments / evaluation:
- ASR tests: Word Error Rate (WER) and Character Error Rate (CER) evaluated on 1,000 utterances with noise levels from −10 dB to +10 dB.
- Objective metrics: PESQ and STOI scores compared for microphone, vibration, and enhanced signals.
- Subjective ratings: 50 participants evaluated audio quality using MUSHRA.
- Real-world tests: Speech recorded in cafés, roadside, trains, and walking scenarios.
- Limitations and future work:
- Whispered speech recognition degrades in very noisy conditions due to weak vibration signals.
- Dynamic input selection (microphone, vibration, or both) based on environmental noise is proposed as future work.
- Physiological variability (e.g., nasal patency) requires per-user calibration and adaptation.
Summary
NasoVoce introduces a novel nose-mounted speech interface combining a microphone and vibration sensor for capturing normal and whispered speech. The system leverages multimodal fusion via the D-DCCRN model to enhance speech quality and recognition accuracy, even in noisy environments. Evaluations confirm its superiority over standalone inputs and commercial alternatives, demonstrating its feasibility for discreet and always-available AI voice interaction. Future work includes dynamic input selection and addressing physiological variability.
Research Questions / Practical Problems
Question signals indexed for this paper.
Based on Jaccard similarity of research subtopics & professions (≥60%)