Aware: Intuitive Device Activation Using Prosody for Natural Voice Interactions
Authors
Document Title
Aware: Intuitive Device Activation Using Prosody for Natural Voice Interactions
Document Information
- Research Area: Human-Computer Interaction, Voice Interaction Technology
- Keywords: Prosody, Intention, Device Activation, Keyword Spotting, Conversational Interface, Voice Interaction
Research Background and Problem
- Problem Description: Current voice interaction devices typically use Keyword Spotting (KWS) to trigger device activation. However, this approach is prone to false activations. For instance, devices may respond to unintentionally mentioned keywords or misrecognize words with similar pronunciations.
- Significance: False activations can lead to privacy breaches and degraded user experience, such as devices unexpectedly starting, recording sensitive information, or performing unintended actions (e.g., accidental purchases). These issues negatively impact the naturalness and reliability of voice interaction devices.
- Research Motivation: Inspired by the role of prosody in conveying meaning during natural language communication, the authors propose leveraging prosodic features to address the false activation problem in voice devices, enabling more natural and intuitive device interaction.
Solution
- Method and Innovation:
- Propose a device activation method based on prosodic features (e.g., pitch variation, intensity, and speech duration) to distinguish between "device-calling" and "non-device-calling" speech patterns.
- Develop a prototype smart speaker named "Aware," which detects prosodic patterns to decide whether to respond to user voice requests, thereby avoiding false activations.
- Utilize machine learning models, such as Convolutional Neural Networks (CNN), to achieve high accuracy in voice classification.
- Implementation Steps and Techniques:
- Data Collection: Gather a voice dataset of device-calling/non-device-calling samples using multiple microphone arrays, including over 90,000 samples covering various directions and distances.
- Speech Feature Analysis: Extract prosodic features (e.g., pitch variation, duration, intensity, HNR) and explore differences between device-calling and non-device-calling speech.
- Classification Model Development: Employ CNNs to classify multidimensional speech features and validate the model's accuracy and robustness.
- Prototype Device Implementation: Build the "Aware" smart speaker integrating voice recognition, feature extraction, and classification models.
- User Experience Evaluation: Conduct user experiments comparing the proposed method with traditional keyword detection methods (e.g., Amazon Echo) to assess activation accuracy and intuitiveness.
Research Results
- Specific Outcomes:
- Data Contribution: Collected and publicly released a dataset of device-calling/non-device-calling voice samples (https://github.com/rkmtlab/Aware) to support future research.
- Feature Identification: Identified prosodic features (e.g., pitch variation curves, speech clarity) as effective indicators of user calling intention.
- Classification Accuracy: The Aware classification model achieved an average accuracy of 86.9% and an F1 score of 0.869 in cross-user testing.
- Prototype Validation: User experiments showed that Aware significantly outperformed Amazon Echo in device response accuracy (F1 score of 0.93 vs. 0.56).
- Advantages:
- Higher calling accuracy, effectively avoiding false activations.
- Requires no additional sensors, relying solely on a single microphone, making it easy to deploy.
- Utilizes natural prosody, resulting in low learning costs and higher user acceptance.
- Experiments and Evaluation:
- User experiments demonstrated that Aware is intuitive and natural to operate, with easy learning and usage.
- Subjective evaluations indicated that participants preferred prosody-based device activation methods over traditional keyword wake-up approaches for daily use.
- Limitations and Future Directions:
- Limitations:
- Calling speech may be influenced by the user's distance from the device and emotional state, leading to variability in prosodic patterns. The model's ability to recognize unseen prosodic patterns is limited.
- Current user experiments were conducted in simulated closed environments. Future validation is needed in real-world scenarios (e.g., offices or homes).
- Future Directions:
- Collect more informal speech data (including calling samples in various contexts and emotional states) to enhance model robustness.
- Explore the application of prosodic patterns in more complex linguistic units (e.g., phrases or sentences) and study free speech or prosody-based device activation methods.
- Develop personalized voice classifiers to further optimize user experience.
- Limitations:
Conclusion
This paper proposes a prosody-based device activation method that effectively addresses the false activation issue in current smart speakers. It validates the potential of prosodic patterns in user intention recognition and provides new research directions for voice interaction design. Future work should focus on expanding the model's generalizability and adapting it to real-world scenarios to promote widespread application of the research outcomes in daily life.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can vocal prosodic features (e.g., pitch variation, intensity, and speech duration) distinguish "device-directed" speech from "non-device-directed" speech?Category: Acoustic Sensing, Activity Recognition, and Speech SecuritySimilar questionsarrow_forward
- Can prosody-based device activation methods significantly improve activation accuracy and intuitiveness of voice devices in real user contexts?Category: Acoustic Sensing, Activity Recognition, and Speech SecuritySimilar questionsarrow_forward
- What significant improvements does prosody-based smart speaker performance offer compared to traditional keyword activation methods?Category: Acoustic Sensing, Activity Recognition, and Speech SecuritySimilar questionsarrow_forward
Practical Problems
1- Voice devices often trigger unexpectedly, causing privacy leaks and degraded user experience.Category: Acoustic Sensing, Activity Recognition, and Speech SecuritySimilar questionsarrow_forward
- 100%
The Impact of User Characteristics and Preferences on Performance with an Unfamiliar Voice User Interface
CHI '19· Voice User Interface (VUI) Design +1
- 100%
ProxiMic: Convenient Voice Activation via Close-to-Mic Speech Detected by a Single Microphone
CHI '21· Voice User Interface (VUI) Design +1
- 100%
Firefox Voice: An Open and Extensible Voice Assistant Built Upon the Web
CHI '21· Voice User Interface (VUI) Design +1
- 100%
Seeking Soulmate via Voice: Understanding Promises and Challenges of Online Synchronized Voice-Based Mobile Dating
CHI '24· Voice User Interface (VUI) Design +1
- 100%
Rewriting the Script: Adapting Text Instructions for Voice Interaction
DIS '23· Voice User Interface (VUI) Design +1
- 100%
Radio2Text: Streaming Speech Recognition Using mmWave Radio Signals
UbiComp '23· Voice User Interface (VUI) Design +1
- 100%
SilentVoice: Unnoticeable Voice Input by Ingressive Speech
UIST '18· Voice User Interface (VUI) Design +1
- 100%
Just Speak It: Minimize Cognitive Load for Text Editing with a Smart Voice Assistant
UIST '21· Voice User Interface (VUI) Design +1
- 100%
App-Based Task Shortcuts for Virtual Assistants
UIST '21· Voice User Interface (VUI) Design +1
- 67%
Comparing Smartphone Speech Recognition and Touchscreen Typing for Composition and Transcription
CHI '20· Voice User Interface (VUI) Design +1
Based on Jaccard similarity of research subtopics & professions (≥60%)