NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech Interaction

Intelligent Voice Assistants (Alexa, Siri, etc.)Affective Human-Computer DialogueContext-Aware ComputingUI/UX DesignersAI/ML Researchers & EngineersSoftware Engineers & Developers

Paper Title

NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech Interaction

Publication Info

  • Topic area: Silent and whispered speech interfaces for AI voice interaction
  • Keywords: silent speech, whispered speech, smart glasses, vibration sensor, microphone, noise robustness, always-available interface, wearable technology, speech recognition, multimodal fusion

Background and Problem

  • Problem / challenge: Existing silent and whispered speech systems fail to balance wearability, silence, noise robustness, and vocabulary size, limiting their practicality for continuous AI interaction.
  • Significance: A robust, unobtrusive, and noise-tolerant speech interface would enable discreet and always-available AI conversations, enhancing usability in public and noisy environments.
  • Motivation and related work: Prior systems like neuromuscular sensors, lip-reading, and bone conduction devices have limitations in vocabulary size, noise robustness, or wearability. Whispered speech systems offer potential but are highly susceptible to noise. This paper addresses these gaps by proposing a novel multimodal approach.

Solution

  • Proposed approach: NasoVoce, a nose-mounted interface integrating a MEMS microphone and a MEMS vibration sensor to capture both air- and bone-conducted speech signals.
  • Novelty:
    1. Integration of a microphone and vibration sensor on the nasal bridge for capturing normal and whispered speech with noise robustness.
    2. Development of a dual-input deep learning model (D-DCCRN) for fusing microphone and vibration sensor signals to enhance speech quality and recognition accuracy.
    3. Creation of a dataset and training methodology for whispered and normal speech in noisy environments.
    4. Demonstration of the system’s practicality through objective metrics, subjective ratings, and real-world evaluations.
  • Procedure and key techniques:
    1. Hardware: A MEMS microphone and vibration sensor mounted on the nose pads of smart glasses, capturing synchronized audio and vibration signals.
    2. Software: A D-DCCRN model processes the real and imaginary components of both signals for noise-robust speech enhancement.
    3. Training: Dataset of 104 hours of paired microphone and vibration sensor recordings with added noise; loss functions include audio enhancement and knowledge distillation.
    4. Evaluation: ASR accuracy, objective quality metrics (PESQ, STOI), subjective ratings (MUSHRA), and real-world tests.

Results

  • Concrete findings:
    • D-DCCRN has 438.25M parameters and processes speech in 136.9 ms, requiring only 31.8% of OpenAI Whisper’s processing time.
    • Enhanced speech outperformed standalone microphone and vibration sensor inputs in ASR accuracy, PESQ, and STOI metrics under most noise conditions.
    • MUSHRA ratings showed enhanced speech consistently scored higher than microphone or vibration-only inputs under various noise levels.
  • Advantage over baselines:
    • Enhanced speech recognition accuracy was superior to microphone-only and vibration-only inputs, particularly in noisy environments.
    • Outperformed Apple AirPods Pro 2’s voice isolation feature in capturing whispered speech.
  • Experiments / evaluation:
    • ASR tests: Word Error Rate (WER) and Character Error Rate (CER) evaluated on 1,000 utterances with noise levels from −10 dB to +10 dB.
    • Objective metrics: PESQ and STOI scores compared for microphone, vibration, and enhanced signals.
    • Subjective ratings: 50 participants evaluated audio quality using MUSHRA.
    • Real-world tests: Speech recorded in cafés, roadside, trains, and walking scenarios.
  • Limitations and future work:
    • Whispered speech recognition degrades in very noisy conditions due to weak vibration signals.
    • Dynamic input selection (microphone, vibration, or both) based on environmental noise is proposed as future work.
    • Physiological variability (e.g., nasal patency) requires per-user calibration and adaptation.

Summary

NasoVoce introduces a novel nose-mounted speech interface combining a microphone and vibration sensor for capturing normal and whispered speech. The system leverages multimodal fusion via the D-DCCRN model to enhance speech quality and recognition accuracy, even in noisy environments. Evaluations confirm its superiority over standalone inputs and commercial alternatives, demonstrating its feasibility for discreet and always-available AI voice interaction. Future work includes dynamic input selection and addressing physiological variability.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222661/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791397
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Intelligent Voice Assistants (Alexa, Siri, etc.), Affective Human-Computer Dialogue, Context-Aware Computing
work
Professions
UI/UX Designers, AI/ML Researchers & Engineers, Software Engineers & Developers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers