DualVoice: Speech Interaction that Discriminates between Normal and Whispered Voice Input

Voice User Interface (VUI) DesignIntelligent Voice Assistants (Alexa, Siri, etc.)Software Engineers & DevelopersUI/UX Designers

Title of the Paper

DualVoice: Speech Interaction that Discriminates between Normal and Whispered Voice Input

Paper Information

  • Domain: Human-Computer Interaction, Speech Recognition
  • Keywords: Speech interaction, Whispered speech recognition, Whispered speech classification, Neural networks, Self-supervised learning, Voice commands, Speech input, Voice privacy

Research Background and Problem

  • Identified Issues or Challenges:

    • Current automatic speech recognition (ASR) systems often confuse commands with text recognition, such as misinterpreting special symbols as regular text.
    • Errors in recognition during editing require manual correction, which disrupts the convenience of hands-free operation.
    • Switching input modes during text input is user-unfriendly and limits the efficiency of speech interaction.
  • Significance:

    • Automatic speech recognition systems are widely used in text input, smart device operation, and navigation systems, but recognition errors severely impact user experience.
    • Providing a simple speech interaction method to distinguish between text and commands can improve efficiency, reduce user confusion, and enhance privacy protection.
  • Research Motivation:

    • To provide a completely hands-free speech interaction solution that simplifies workflows.
    • To improve speech recognition using standard microphones without requiring additional hardware investment.

Solution

  • Method or Solution:

    • Introduced DualVoice, a speech interaction method that uses whispered input for commands and normal voice for text input.
    • Developed two types of neural networks: one for distinguishing between normal and whispered speech, and another for recognizing whispered speech.
  • Innovations:

    • Leveraged whispered speech as a means of command input, enabling the distinction between text and commands through natural switching between normal and whispered speech.
    • Addressed privacy and social acceptability issues of speech interaction in public spaces.
    • Achieved functionality using only standard microphones, eliminating the need for additional hardware.
  • Implementation Steps and Key Techniques:

    1. Pretrained models using self-supervised learning (e.g., wav2vec 2.0 and HuBERT neural networks) to learn features from unlabeled speech data.
    2. Fine-tuned the models with a small amount of labeled data to enhance whispered speech recognition capabilities.
    3. Classified speech streams into normal and whispered categories, directing them to the appropriate speech recognizer.
    4. Designed user interface prototypes to demonstrate practical applications of whispered and normal speech interaction.

Research Findings

  • Specific Outcomes:

    • Developed a speech interface capable of distinguishing between normal and whispered speech and executing corresponding actions.
    • Built a whispered speech recognizer based on self-supervised learning, achieving a high-accuracy whispered speech classification system.
    • Demonstrated the effectiveness of the method through visually intuitive prototype testing.
  • Advantages:

    • Command input via whispered speech is intuitive, natural, and socially acceptable.
    • Does not require specialized hardware, making it cost-effective and fully supportive of hands-free interaction.
    • Achieved high classification and recognition accuracy in experiments (whispered speech classification accuracy reached 96.7%).
  • Experimental or Evaluation Results:

    • The trained neural networks effectively recognized whispered speech after fine-tuning (minimum word error rate of 0.38%).
    • The whispered speech classification model performed consistently on wTIMIT training data and user-customized data, demonstrating strong scalability.
    • User experiments showed that the system has a low learning curve, and users found it easy to switch between whispered and normal speech.
  • Limitations and Future Directions:

    • Current whispered speech recognition relies on user-customized training and cannot yet achieve fully independent user recognition.
    • Further exploration is needed for additional applications of whispered speech recognition, such as remote meeting command input, multi-character control in virtual environments, and integration with silent speech technologies to expand interaction modes.

Additional Information (if applicable)

  • This research was proposed based on extensive literature review and serves as a comparison and complement to existing audio interaction technologies (e.g., Google Cloud Speech-to-Text, SilentVoice).
  • The first author outlined clear research directions for future exploration of whispered and silent speech, as well as the development of public whispered speech datasets.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/uist/85048/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3526113.3545685
At a Glance

Paper Snapshot

fact_check
dataset
Source
UIST
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
1 authors
sell
Subtopics
Voice User Interface (VUI) Design, Intelligent Voice Assistants (Alexa, Siri, etc.)
work
Professions
Software Engineers & Developers, UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers