mSilent: Towards General Corpus Silent Speech Recognition Using COTS mmWave Radar
Authors
"Silent speech recognition (SSR) allows users to speak to the device without making a sound, avoiding being overheard or disturbing others. Compared to the video-based approach, wireless signal-based SSR can work when the user is wearing a mask and has fewer privacy concerns. However, previous wireless-based systems are still far from well-studied, e.g. they are only evaluated in corpus with highly limited size, making them only feasible for interaction with dozens of deterministic commands. In this paper, we present mSilent, a millimeter-wave (mmWave) based SSR system that can work in the general corpus containing thousands of daily conversation sentences. With the strong recognition capability, mSilent not only supports the more complex interaction with assistants, but also enables more general applications in daily life such as communication and input. To extract fine-grained articulatory features, we build a signal processing pipeline that uses a clustering-selection algorithm to separate articulatory gestures and generates a multi-scale detrended spectrogram (MSDS). To handle the complexity of the general corpus, we design an end-to-end deep neural network that consists of a multi-branch convolutional front-end and a Transformer-based sequence-to-sequence back-end. We collect a general corpus dataset of 1,000 daily conversation sentences that contains 21K samples of bi-modality data (mmWave and video). Our evaluation shows that mSilent achieves a 9.5% average word error rate (WER) at a distance of 1.5m, which is comparable to the performance of the state-of-the-art video-based approach. We also explore deploying mSilent in two typical scenarios of text entry and in-car assistant, and the less than 6% average WER demonstrates the potential of mSilent in general daily applications. https://dl.acm.org/doi/10.1145/3580838"
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can COTS mmWave radar enable silent speech recognition for general vocabulary?Category: Silent Speech, Whisper, and Lip-Movement InteractionSimilar questionsarrow_forward
- How can signal processing and deep learning methods be designed to capture fine vocal cord and mouth movement features?Category: Silent Speech, Whisper, and Lip-Movement InteractionSimilar questionsarrow_forward
- How can user-adaptive learning address inter-individual differences to improve recognition accuracy?Category: Silent Speech, Whisper, and Lip-Movement InteractionSimilar questionsarrow_forward
Practical Problems
1- Users struggle to achieve privacy-friendly voice interaction in noisy environments or while wearing masks.Category: Silent Speech, Whisper, and Lip-Movement InteractionSimilar questionsarrow_forward
- 67%
The Impact of User Characteristics and Preferences on Performance with an Unfamiliar Voice User Interface
CHI '19· Voice User Interface (VUI) Design +1
- 67%
ProxiMic: Convenient Voice Activation via Close-to-Mic Speech Detected by a Single Microphone
CHI '21· Voice User Interface (VUI) Design +1
- 67%
Firefox Voice: An Open and Extensible Voice Assistant Built Upon the Web
CHI '21· Voice User Interface (VUI) Design +1
- 67%
Aware: Intuitive Device Activation Using Prosody for Natural Voice Interactions
CHI '22· Voice User Interface (VUI) Design +1
- 67%
Seeking Soulmate via Voice: Understanding Promises and Challenges of Online Synchronized Voice-Based Mobile Dating
CHI '24· Voice User Interface (VUI) Design +1
- 67%
Cross-Cultural Validation of Partner Models for Voice User Interfaces
CUI '24· Voice User Interface (VUI) Design +1
- 67%
Rewriting the Script: Adapting Text Instructions for Voice Interaction
DIS '23· Voice User Interface (VUI) Design +1
- 67%
Radio2Text: Streaming Speech Recognition Using mmWave Radio Signals
UbiComp '23· Voice User Interface (VUI) Design +1
- 67%
SilentVoice: Unnoticeable Voice Input by Ingressive Speech
UIST '18· Voice User Interface (VUI) Design +1
- 67%
Just Speak It: Minimize Cognitive Load for Text Editing with a Smart Voice Assistant
UIST '21· Voice User Interface (VUI) Design +1
Based on Jaccard similarity of research subtopics & professions (≥60%)