Research Background and Issues

  • Identified Problems or Challenges: The authors found that current voice interaction systems face issues such as insufficient privacy protection, severe noise interference, and lack of support for silent voice interaction in shared spaces. In scenarios like museums, vehicles, and offices, users may be reluctant to speak loudly or have privacy concerns, limiting the application of voice control systems. Additionally, existing silent voice interfaces often require auxiliary devices such as cameras, millimeter-wave radars, or wearable devices, which pose inconvenience and high costs.
  • Importance: With the widespread adoption of voice-controlled devices, privacy concerns are becoming increasingly prominent. Addressing these issues can not only expand the application scenarios of voice interaction but also enhance user experience and social acceptance.
  • Research Motivation and Related Work: Although prior research in the field has explored various methods (e.g., silent voice recognition based on electromyography, camera detection, or millimeter-wave radar), these approaches are typically limited to single-user setups or fail to meet privacy and convenience requirements. The authors aim to develop a solution that supports multi-user interaction in shared environments without requiring wearable devices.

Solution

  • Proposed Solution: The authors introduced the M2Silent system, which integrates multi-directional speakers, frequency-modulated continuous wave (FMCW) signals, and deep learning technologies to enable simultaneous silent voice interaction for multiple users. The system combines audio transmission with silent voice sensing capabilities.
  • Innovations:
    1. For the first time, multi-directional speakers are combined with silent voice interaction, enabling device-free interaction for multiple users simultaneously.
    2. FMCW signals are used as audio carriers to simultaneously facilitate voice playback and user surround sensing.
    3. A time-shifted FMCW signal and blind signal separation algorithm are proposed to extract silent voice features from multiple users.
    4. A sliding window approach and the deep residual model SilentMatch are employed for real-time and sentence-level silent voice recognition.
  • Implementation Steps:
    1. Modulate audio into FMCW signals and utilize air nonlinearity for demodulation to achieve clear audio playback.
    2. Distribute signals to different directions via time shifting to differentiate user features, applying blind signal separation techniques to isolate user characteristics.
    3. Use the SilentMatch deep learning model for word-level and sentence-level voice recognition.
    4. Employ a sliding window method in real-time interaction to process word sequences, ensuring smooth interaction.

Research Outcomes

  • Specific Results:
    1. Achieved device-free, simultaneous silent voice interaction for multiple users.
    2. In real-world testing, the M2Silent system achieved a word error rate (WER) of 6.5%, a sequence error rate (SER) of 12.8%, and excellent audio quality (PESQ score of 2.81).
  • Advantages Over Existing Solutions:
    1. Compared to wearable device-based solutions (e.g., headphones or smart glasses), it reduces user burden and costs.
    2. Compared to vision-based methods, it addresses privacy concerns and limitations due to ambient lighting while maintaining high recognition accuracy.
    3. Supports simultaneous interaction for multiple users, which existing single-user silent voice systems cannot achieve.
  • Experimental or Evaluation Results: Extensive experiments tested the system's performance, analyzing the impact of key parameters (e.g., signal shape, number of users, direction, and distance) on voice quality and recognition accuracy. Results demonstrated stable performance in multi-user scenarios, low interaction latency, and moderate resource costs.
  • Limitations and Future Directions:
    1. The maximum supported user count is three; additional users may lead to signal overlap and reduced recognition performance.
    2. The current maximum effective distance is approximately 2 meters; long-distance applications require improvements in sound wave emission power or deployment of additional devices.
    3. Obstructions can affect sound propagation and reflected signals; solutions such as leveraging environmental reflections or distributed multi-device systems need exploration.
    4. Regarding health and safety, directional ultrasound must comply with international standards. Future optimization of signal frequency could support longer distances and higher energy levels.

Conclusion

The M2Silent system effectively achieves efficient, privacy-preserving multi-user silent voice interaction by integrating multi-directional speakers, FMCW signals, and deep learning technologies. This innovation not only enhances the practicality of the technology but also expands the application scenarios of voice interaction, particularly for vehicles, museums, and public spaces. However, further improvements are needed to increase the system's user capacity, operational range, and adaptability to complex environments, thereby unlocking its full application potential.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189260/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714174
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
8 authors
sell
Subtopics
Eye Tracking & Gaze Interaction, Voice User Interface (VUI) Design, Privacy by Design & User Control
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
8 related papers