DualVoice: Speech Interaction that Discriminates between Normal and Whispered Voice Input
Authors
Title of the Paper
DualVoice: Speech Interaction that Discriminates between Normal and Whispered Voice Input
Paper Information
- Domain: Human-Computer Interaction, Speech Recognition
- Keywords: Speech interaction, Whispered speech recognition, Whispered speech classification, Neural networks, Self-supervised learning, Voice commands, Speech input, Voice privacy
Research Background and Problem
-
Identified Issues or Challenges:
- Current automatic speech recognition (ASR) systems often confuse commands with text recognition, such as misinterpreting special symbols as regular text.
- Errors in recognition during editing require manual correction, which disrupts the convenience of hands-free operation.
- Switching input modes during text input is user-unfriendly and limits the efficiency of speech interaction.
-
Significance:
- Automatic speech recognition systems are widely used in text input, smart device operation, and navigation systems, but recognition errors severely impact user experience.
- Providing a simple speech interaction method to distinguish between text and commands can improve efficiency, reduce user confusion, and enhance privacy protection.
-
Research Motivation:
- To provide a completely hands-free speech interaction solution that simplifies workflows.
- To improve speech recognition using standard microphones without requiring additional hardware investment.
Solution
-
Method or Solution:
- Introduced DualVoice, a speech interaction method that uses whispered input for commands and normal voice for text input.
- Developed two types of neural networks: one for distinguishing between normal and whispered speech, and another for recognizing whispered speech.
-
Innovations:
- Leveraged whispered speech as a means of command input, enabling the distinction between text and commands through natural switching between normal and whispered speech.
- Addressed privacy and social acceptability issues of speech interaction in public spaces.
- Achieved functionality using only standard microphones, eliminating the need for additional hardware.
-
Implementation Steps and Key Techniques:
- Pretrained models using self-supervised learning (e.g., wav2vec 2.0 and HuBERT neural networks) to learn features from unlabeled speech data.
- Fine-tuned the models with a small amount of labeled data to enhance whispered speech recognition capabilities.
- Classified speech streams into normal and whispered categories, directing them to the appropriate speech recognizer.
- Designed user interface prototypes to demonstrate practical applications of whispered and normal speech interaction.
Research Findings
-
Specific Outcomes:
- Developed a speech interface capable of distinguishing between normal and whispered speech and executing corresponding actions.
- Built a whispered speech recognizer based on self-supervised learning, achieving a high-accuracy whispered speech classification system.
- Demonstrated the effectiveness of the method through visually intuitive prototype testing.
-
Advantages:
- Command input via whispered speech is intuitive, natural, and socially acceptable.
- Does not require specialized hardware, making it cost-effective and fully supportive of hands-free interaction.
- Achieved high classification and recognition accuracy in experiments (whispered speech classification accuracy reached 96.7%).
-
Experimental or Evaluation Results:
- The trained neural networks effectively recognized whispered speech after fine-tuning (minimum word error rate of 0.38%).
- The whispered speech classification model performed consistently on wTIMIT training data and user-customized data, demonstrating strong scalability.
- User experiments showed that the system has a low learning curve, and users found it easy to switch between whispered and normal speech.
-
Limitations and Future Directions:
- Current whispered speech recognition relies on user-customized training and cannot yet achieve fully independent user recognition.
- Further exploration is needed for additional applications of whispered speech recognition, such as remote meeting command input, multi-character control in virtual environments, and integration with silent speech technologies to expand interaction modes.
Additional Information (if applicable)
- This research was proposed based on extensive literature review and serves as a comparison and complement to existing audio interaction technologies (e.g., Google Cloud Speech-to-Text, SilentVoice).
- The first author outlined clear research directions for future exploration of whispered and silent speech, as well as the development of public whispered speech datasets.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can whisper input and normal speech input be distinguished for commands versus text to improve voice interaction efficiency?Category: Silent Speech, Whisper, and Lip-Movement InteractionSimilar questionsarrow_forward
- How can whisper speech recognition with standard microphones achieve high-accuracy classification and recognition?Category: Silent Speech, Whisper, and Lip-Movement InteractionSimilar questionsarrow_forward
- Can whisper input improve privacy protection and social acceptability of voice interaction?Category: Silent Speech, Whisper, and Lip-Movement InteractionSimilar questionsarrow_forward
Practical Problems
1- Speech recognition systems easily confuse commands and text, and manual correction is tedious and disrupts workflow.Category: Silent Speech, Whisper, and Lip-Movement InteractionSimilar questionsarrow_forward
- 100%
Voicify Your UI: Towards Android App Control with Voice Commands
UbiComp '23· Voice User Interface (VUI) Design +1
- 80%
Panel: Voice Assistants, UX Design and Research
CHI '18· Voice User Interface (VUI) Design +1
- 75%
Keep it Short: A Comparison of Voice Assistants' Response Behavior
CHI '22· Voice User Interface (VUI) Design +1
- 75%
VOICON: Geometric Motion-Based Visual Feedback in Voice User Interface
DIS '24· Voice User Interface (VUI) Design +1
- 60%
Speech and Hands-free Interaction: Myths, Challenges, and Opportunities
CHI '18· Voice User Interface (VUI) Design +1
- 60%
The Impact of Word, Multiple Word, and Sentence Input on Virtual Keyboard Decoding Performance
CHI '18· Voice User Interface (VUI) Design +1
- 60%
EYEditor: Towards On-the-Go Heads-Up Text Editing Using Voice and Manual Input
CHI '20· Head-Up Display (HUD) & Advanced Driver Assistance Systems (ADAS) +1
- 60%
Collecting and Characterizing Natural Language Utterances for Specifying Data Visualizations
CHI '21· Voice User Interface (VUI) Design +1
- 60%
EmoWear: Exploring Emotional Teasers for Voice Message Interaction on Smartwatches
CHI '24· Haptic Wearables +2
- 60%
PonDeFlick: A Japanese Text Entry on Smartwatch Commonalizing Flick Operation with Smartphone Interface
CHI '24· Foot & Wrist Interaction +1
Based on Jaccard similarity of research subtopics & professions (≥60%)