ProxiMic: Convenient Voice Activation via Close-to-Mic Speech Detected by a Single Microphone

Voice User Interface (VUI) DesignIntelligent Voice Assistants (Alexa, Siri, etc.)

Title of the Paper

ProxiMic: Convenient Voice Activation via Close-to-Mic Speech Detected by a Single Microphone

Paper Information

  • Research Area: Human-Computer Interaction, Voice Activation Technology
  • Keywords: Voice Input, Sensing Technology, Activity Recognition, Close-to-Mic Speech, Privacy Protection, User Experience, Machine Learning, Human-Computer Interaction, Perception Technology, Adverse Environment Detection

Research Background and Problem

  • Challenges:
    • Voice input faces three main challenges: prolonged wake-up time, cumbersome and frustrating multi-turn interaction processes, and risks of voice privacy leakage.
    • Existing methods such as "Raise to Speak" and "PrivateTalk" have made breakthroughs in reducing wake-up time, but they rely on multiple sensors or specific hardware configurations, limiting their applicability.
    • Traditional privacy protection methods like silent speech are effective but involve complex devices and support limited semantic ranges, preventing widespread adoption.
  • Importance of the Problem:
    • With the increasing prevalence of voice input technology, addressing voice privacy and interaction experience issues is crucial for enhancing user experience.
  • Research Motivation and Related Work:
    • To provide a wake-free, low-cost voice input method that addresses wake words, cumbersome interactions, and privacy concerns.
    • The proposed technology can adapt to various device formats without requiring additional sensors or specialized equipment.

Solution

  • Method or Solution:
    • A technology named ProxiMic is proposed, enabling wake-free activation via detecting "close-to-mic speech" using a single microphone. It supports users in conducting direct voice interactions across multiple conversational turns.
    • CNN (Convolutional Neural Network) is employed to detect "pop noise" and other subtle close-to-mic features.
    • A two-stage algorithm is designed:
      • The first stage uses an Adaptive Amplitude Threshold Trigger (AATT) to filter potential close-to-mic speech signals.
      • The second stage employs CNN to determine whether the signal is close-to-mic speech for further precise identification.
  • Innovations:
    • Innovative use of pop noise as a sound feature for voice activation without requiring multi-sensor support.
    • Development and optimization of a two-stage detection algorithm, balancing high accuracy, low power consumption, and low memory usage.
  • Implementation Steps and Key Technologies:
    • AATT: Dynamically monitors environmental background sound energy to filter high-amplitude potential signals.
    • CNN: Refines and distinguishes close-to-mic speech signals based on sound spectrogram features (low-frequency enhanced spectrogram).
    • Experiments simulate real-world scenarios with various device environments and validate robustness against complex background noise.

Research Results

  • Specific Results:
    • Achieved a 94.1% activation recall rate and 12.3 false accepts per week per user.
    • Demonstrated adaptability of ProxiMic across various device formats (e.g., smartphones, watches, headphones) with high environmental robustness.
  • Advantages:
    • Does not require special gestures or additional sensors, making it low-cost and easy to deploy.
    • Significantly enhances voice input privacy, allowing users to issue voice commands with low volume or whispering.
    • Outperforms existing algorithms (e.g., "Raise to Speak") in memory usage and processing efficiency.
  • Experimental or Evaluation Results:
    • User experiments indicate that compared to traditional wake words or button activation methods, ProxiMic achieves higher efficiency ratings and user satisfaction.
    • ASR experimental results show high transcription accuracy even with pop noise or low-volume whispering.
  • Limitations and Future Directions:
    • Multi-device adaptation requires further data training and testing to support various hardware configurations.
    • Optimizing semantic recognition for whispering and further reducing false trigger rates are key future research directions.
    • Enhancing robustness in more complex linguistic environments and adverse acoustic scenarios remains a challenge.

Conclusion

This study proposes an innovative and efficient close-to-mic voice activation technology, addressing the interaction complexity and privacy issues inherent in traditional wake-up and voice input methods. Experimental results demonstrate that ProxiMic significantly improves user experience and has broad application potential, such as privacy-sensitive scenarios and IoT device voice interactions. The research provides insights for the future development of voice input technology while highlighting challenges and opportunities in multi-device adaptation and feature expansion.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/47424/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3411764.3445687
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Voice User Interface (VUI) Design, Intelligent Voice Assistants (Alexa, Siri, etc.)
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
10 related papers