Aware: Intuitive Device Activation Using Prosody for Natural Voice Interactions

Voice User Interface (VUI) DesignIntelligent Voice Assistants (Alexa, Siri, etc.)

Document Title

Aware: Intuitive Device Activation Using Prosody for Natural Voice Interactions

Document Information

  • Research Area: Human-Computer Interaction, Voice Interaction Technology
  • Keywords: Prosody, Intention, Device Activation, Keyword Spotting, Conversational Interface, Voice Interaction

Research Background and Problem

  • Problem Description: Current voice interaction devices typically use Keyword Spotting (KWS) to trigger device activation. However, this approach is prone to false activations. For instance, devices may respond to unintentionally mentioned keywords or misrecognize words with similar pronunciations.
  • Significance: False activations can lead to privacy breaches and degraded user experience, such as devices unexpectedly starting, recording sensitive information, or performing unintended actions (e.g., accidental purchases). These issues negatively impact the naturalness and reliability of voice interaction devices.
  • Research Motivation: Inspired by the role of prosody in conveying meaning during natural language communication, the authors propose leveraging prosodic features to address the false activation problem in voice devices, enabling more natural and intuitive device interaction.

Solution

  • Method and Innovation:
    • Propose a device activation method based on prosodic features (e.g., pitch variation, intensity, and speech duration) to distinguish between "device-calling" and "non-device-calling" speech patterns.
    • Develop a prototype smart speaker named "Aware," which detects prosodic patterns to decide whether to respond to user voice requests, thereby avoiding false activations.
    • Utilize machine learning models, such as Convolutional Neural Networks (CNN), to achieve high accuracy in voice classification.
  • Implementation Steps and Techniques:
    1. Data Collection: Gather a voice dataset of device-calling/non-device-calling samples using multiple microphone arrays, including over 90,000 samples covering various directions and distances.
    2. Speech Feature Analysis: Extract prosodic features (e.g., pitch variation, duration, intensity, HNR) and explore differences between device-calling and non-device-calling speech.
    3. Classification Model Development: Employ CNNs to classify multidimensional speech features and validate the model's accuracy and robustness.
    4. Prototype Device Implementation: Build the "Aware" smart speaker integrating voice recognition, feature extraction, and classification models.
    5. User Experience Evaluation: Conduct user experiments comparing the proposed method with traditional keyword detection methods (e.g., Amazon Echo) to assess activation accuracy and intuitiveness.

Research Results

  • Specific Outcomes:
    • Data Contribution: Collected and publicly released a dataset of device-calling/non-device-calling voice samples (https://github.com/rkmtlab/Aware) to support future research.
    • Feature Identification: Identified prosodic features (e.g., pitch variation curves, speech clarity) as effective indicators of user calling intention.
    • Classification Accuracy: The Aware classification model achieved an average accuracy of 86.9% and an F1 score of 0.869 in cross-user testing.
    • Prototype Validation: User experiments showed that Aware significantly outperformed Amazon Echo in device response accuracy (F1 score of 0.93 vs. 0.56).
  • Advantages:
    • Higher calling accuracy, effectively avoiding false activations.
    • Requires no additional sensors, relying solely on a single microphone, making it easy to deploy.
    • Utilizes natural prosody, resulting in low learning costs and higher user acceptance.
  • Experiments and Evaluation:
    • User experiments demonstrated that Aware is intuitive and natural to operate, with easy learning and usage.
    • Subjective evaluations indicated that participants preferred prosody-based device activation methods over traditional keyword wake-up approaches for daily use.
  • Limitations and Future Directions:
    • Limitations:
      • Calling speech may be influenced by the user's distance from the device and emotional state, leading to variability in prosodic patterns. The model's ability to recognize unseen prosodic patterns is limited.
      • Current user experiments were conducted in simulated closed environments. Future validation is needed in real-world scenarios (e.g., offices or homes).
    • Future Directions:
      • Collect more informal speech data (including calling samples in various contexts and emotional states) to enhance model robustness.
      • Explore the application of prosodic patterns in more complex linguistic units (e.g., phrases or sentences) and study free speech or prosody-based device activation methods.
      • Develop personalized voice classifiers to further optimize user experience.

Conclusion

This paper proposes a prosody-based device activation method that effectively addresses the false activation issue in current smart speakers. It validates the potential of prosodic patterns in user intention recognition and provides new research directions for voice interaction design. Future work should focus on expanding the model's generalizability and adapting it to real-world scenarios to promote widespread application of the research outcomes in daily life.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/71872/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/abs/10.1145/3491102.3517687
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Voice User Interface (VUI) Design, Intelligent Voice Assistants (Alexa, Siri, etc.)
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
10 related papers