M^2Silent: Enabling Multi-user Silent Speech Interactions via Multi-directional Speakers in Shared Spaces
Authors
Research Background and Issues
- Identified Problems or Challenges: The authors found that current voice interaction systems face issues such as insufficient privacy protection, severe noise interference, and lack of support for silent voice interaction in shared spaces. In scenarios like museums, vehicles, and offices, users may be reluctant to speak loudly or have privacy concerns, limiting the application of voice control systems. Additionally, existing silent voice interfaces often require auxiliary devices such as cameras, millimeter-wave radars, or wearable devices, which pose inconvenience and high costs.
- Importance: With the widespread adoption of voice-controlled devices, privacy concerns are becoming increasingly prominent. Addressing these issues can not only expand the application scenarios of voice interaction but also enhance user experience and social acceptance.
- Research Motivation and Related Work: Although prior research in the field has explored various methods (e.g., silent voice recognition based on electromyography, camera detection, or millimeter-wave radar), these approaches are typically limited to single-user setups or fail to meet privacy and convenience requirements. The authors aim to develop a solution that supports multi-user interaction in shared environments without requiring wearable devices.
Solution
- Proposed Solution: The authors introduced the M2Silent system, which integrates multi-directional speakers, frequency-modulated continuous wave (FMCW) signals, and deep learning technologies to enable simultaneous silent voice interaction for multiple users. The system combines audio transmission with silent voice sensing capabilities.
- Innovations:
- For the first time, multi-directional speakers are combined with silent voice interaction, enabling device-free interaction for multiple users simultaneously.
- FMCW signals are used as audio carriers to simultaneously facilitate voice playback and user surround sensing.
- A time-shifted FMCW signal and blind signal separation algorithm are proposed to extract silent voice features from multiple users.
- A sliding window approach and the deep residual model SilentMatch are employed for real-time and sentence-level silent voice recognition.
- Implementation Steps:
- Modulate audio into FMCW signals and utilize air nonlinearity for demodulation to achieve clear audio playback.
- Distribute signals to different directions via time shifting to differentiate user features, applying blind signal separation techniques to isolate user characteristics.
- Use the SilentMatch deep learning model for word-level and sentence-level voice recognition.
- Employ a sliding window method in real-time interaction to process word sequences, ensuring smooth interaction.
Research Outcomes
- Specific Results:
- Achieved device-free, simultaneous silent voice interaction for multiple users.
- In real-world testing, the M2Silent system achieved a word error rate (WER) of 6.5%, a sequence error rate (SER) of 12.8%, and excellent audio quality (PESQ score of 2.81).
- Advantages Over Existing Solutions:
- Compared to wearable device-based solutions (e.g., headphones or smart glasses), it reduces user burden and costs.
- Compared to vision-based methods, it addresses privacy concerns and limitations due to ambient lighting while maintaining high recognition accuracy.
- Supports simultaneous interaction for multiple users, which existing single-user silent voice systems cannot achieve.
- Experimental or Evaluation Results: Extensive experiments tested the system's performance, analyzing the impact of key parameters (e.g., signal shape, number of users, direction, and distance) on voice quality and recognition accuracy. Results demonstrated stable performance in multi-user scenarios, low interaction latency, and moderate resource costs.
- Limitations and Future Directions:
- The maximum supported user count is three; additional users may lead to signal overlap and reduced recognition performance.
- The current maximum effective distance is approximately 2 meters; long-distance applications require improvements in sound wave emission power or deployment of additional devices.
- Obstructions can affect sound propagation and reflected signals; solutions such as leveraging environmental reflections or distributed multi-device systems need exploration.
- Regarding health and safety, directional ultrasound must comply with international standards. Future optimization of signal frequency could support longer distances and higher energy levels.
Conclusion
The M2Silent system effectively achieves efficient, privacy-preserving multi-user silent voice interaction by integrating multi-directional speakers, FMCW signals, and deep learning technologies. This innovation not only enhances the practicality of the technology but also expands the application scenarios of voice interaction, particularly for vehicles, museums, and public spaces. However, further improvements are needed to increase the system's user capacity, operational range, and adaptability to complex environments, thereby unlocking its full application potential.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can silent voice interaction protect privacy and reduce noise interference?Category: Silent Speech, Whisper, and Lip-Movement InteractionSimilar questionsarrow_forward
- In multi-user shared spaces, how can efficient silent speech recognition be achieved without device support?Category: Silent Speech, Whisper, and Lip-Movement InteractionSimilar questionsarrow_forward
- How can multidirectional speakers and deep learning be integrated to improve silent voice interaction accuracy and practicality?Category: Silent Speech, Whisper, and Lip-Movement InteractionSimilar questionsarrow_forward
Practical Problems
1- In shared spaces, users avoid speaking aloud due to privacy concerns, limiting voice control use.Category: Silent Speech, Whisper, and Lip-Movement InteractionSimilar questionsarrow_forward
- 67%
An Evaluation of Radar Metaphors for Providing Directional Stimuli Using Non-Verbal Sound
CHI '19· Eye Tracking & Gaze Interaction +1
- 67%
TAGSwipe: Touch Assisted Gaze Swipe for Text Entry
CHI '20· Eye Tracking & Gaze Interaction +1
- 67%
Hummer: Text Entry by Gaze and Hum
CHI '21· Eye Tracking & Gaze Interaction +1
- 67%
Integrating Gaze and Speech for Enabling Implicit Interactions
CHI '22· Eye Tracking & Gaze Interaction +1
- 67%
Investigating the Effects of Simulated Eye Contact in Video Call Interviews
CHI '25· Eye Tracking & Gaze Interaction +1
- 67%
EyeSayCorrect: Eye Gaze and Voice Based Hands-free Text Correction for Mobile Devices
IUI '22· Eye Tracking & Gaze Interaction +1
- 67%
PrivateGaze: Preserving User Privacy in Black-box Mobile Gaze Tracking Services
UbiComp '24· Eye Tracking & Gaze Interaction +1
- 60%
PrivateTalk: Activating Voice Input with Hand-On-Mouth Gesture Detected by Bluetooth Earphones
UIST '19· Eye Tracking & Gaze Interaction +2
Based on Jaccard similarity of research subtopics & professions (≥60%)