EchoSpeech: Continuous Silent Speech Recognition on Minimally-obtrusive Eyewear Powered by Acoustic Sensing

Vibrotactile Feedback & Skin StimulationVoice User Interface (VUI) DesignBiosensors & Physiological MonitoringAI/ML Researchers & EngineersAssistive Technology Specialists

Document Title

EchoSpeech: Continuous Silent Speech Recognition on Minimally-obtrusive Eyewear Powered by Acoustic Sensing

Document Information

  • Subject Area: Human-Computer Interaction, Speech Recognition, Silent Speech Interfaces
  • Keywords: Silent Speech Recognition, Acoustic Sensing, Smart Glasses, Continuous Speech Recognition, Deep Learning, Miniature Sensors, Low Power, Human-Computer Interface

Research Background and Problem

  • Problems and Challenges:

    • Silent Speech Interfaces (SSI) enable users to interact without vocalizing, making them crucial for noisy environments or situations requiring silence.
    • Current silent speech recognition technologies face several limitations. Camera-based methods require a clear frontal view and raise privacy concerns, while contact-based sensors (e.g., those attached to the mouth or chin) cause physical discomfort and are inconvenient for prolonged use.
    • Existing devices are often limited to recognizing a small set of command words and lack the ability to handle continuous speech effectively.
  • Significance:

    • Silent speech interfaces can expand the application scenarios of voice assistants, such as entering passwords or performing silent controls in shared spaces.
    • Providing a wearable and minimally intrusive solution can address the physical and social comfort issues present in existing technologies.
  • Motivation and Related Work:

    • Previous studies have explored devices placed behind the ear, on the back of the chin, or inside the ear. However, these approaches suffer from limited signal quality. Camera-based methods face challenges in portability and privacy, making acoustic sensing a promising research direction.
    • Early studies like SpeeChin and EarCommand demonstrated some capabilities but fell short in terms of speech recognition speed, cross-session performance, and natural-style continuous speech recognition.

Solution

  • Method and Solution:

    • A solution named EchoSpeech is proposed, which achieves contactless, low-power silent speech recognition by integrating acoustic sensors (speakers and microphones) into commercial off-the-shelf glass frames.
    • A custom deep learning pipeline is designed, leveraging a Connectionist Temporal Classification (CTC) loss function to enable continuous and discrete speech recognition without manual segmentation of speech.
    • A two-step training approach ("pre-training + fine-tuning") is introduced to minimize the amount of user data required.
  • Innovations:

    1. Hardware Innovation: The first implementation of a contactless silent speech interface on a single glass frame.
    2. Algorithmic Innovation: A segmentation-independent speech recognition pipeline based on Convolutional Neural Networks (CNNs).
    3. Robust Design: Evaluated in complex scenarios, including walking and noise injection.
  • Implementation Steps and Technical Highlights:

    1. Hardware Design:
      • Two sets of speakers and microphones are installed on both sides of the glass frame, forming four primary sound reflection paths.
      • Frequency-Modulated Continuous Wave (FMCW) signals and differential echo analysis are used to capture facial motion patterns.
    2. Deep Learning Model:
      • ResNet-18 is employed as the CNN encoder.
      • The CTC loss function addresses the temporal alignment of variable-length sequences.
    3. Experiments and Optimization:
      • Data augmentation (e.g., random noise, simulated moving scene noise) is applied to enhance model robustness.
      • User-specific fine-tuning reduces training time and improves accuracy.

Research Outcomes

  • Specific Results:

    • Achieved an average Word Error Rate (WER) of 4.5% (standard deviation 3.5%) for 31 isolated command words and a WER of 6.1% (standard deviation 4.2%) for 3-6 digit continuous input tasks.
    • Demonstrated strong robustness in walking and complex noise environments.
  • Comparison with Existing Solutions:

    • Compared to existing SSI methods (e.g., camera-based, behind-the-ear sensors, or adhesive sensors), EchoSpeech offers higher comfort and avoids direct skin contact, making it more socially acceptable.
    • In terms of power consumption, the system operates at 73.3mW, significantly lower than camera-based methods (e.g., SpeeChin at 2.4W).
  • Experimental and Evaluation Results:

    • Achieved segmentation-independent continuous prediction using a sliding window mechanism.
    • In walking environments (without additional training), the system showed a WER of 16.8%; after data augmentation and fine-tuning with a small amount of training data (6-8 minutes), the WER was reduced to 8.7%.
    • The system maintained stable performance across varying input speeds and sequence lengths.
  • Limitations and Future Directions:

    • Limitations:
      1. Lack of noise resistance for operations like adjusting the glasses.
      2. Instability in wearing for users with certain facial shapes.
      3. Further optimization is needed for power consumption and environmental noise handling.
    • Future Directions:
      1. Expand data collection to improve cross-user generalization.
      2. Integrate more flexible activation mechanisms to save power.
      3. Further optimize the device's inconspicuous design to enhance social acceptability.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/96520/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3580801
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Vibrotactile Feedback & Skin Stimulation, Voice User Interface (VUI) Design, Biosensors & Physiological Monitoring
work
Professions
AI/ML Researchers & Engineers, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
0 related papers