WESPER: Zero-shot and Realtime Whisper to Normal Voice Conversion for Whisper-based Speech interactions

Intelligent Voice Assistants (Alexa, Siri, etc.)Voice AccessibilitySpeech-Language Pathologists & Audiologists

Document Title

WESPER: Zero-shot and Realtime Whisper to Normal Voice Conversion for Whisper-based Speech Interactions

Document Information

  • Subject Area: Human-computer speech interaction, voice conversion, assistive technologies for hearing and speech impairments
  • Keywords: speech interaction, whisper voice, whisper conversion, artificial intelligence, speech recognition, self-supervised learning, speech reconstruction

Research Background and Problem

  • Identified Problems or Challenges:

    • Using voice commands in public spaces may lead to privacy concerns or social discomfort.
    • In scenarios like conference calls, there may be risks of environmental disturbance and information leakage.
    • For individuals with speech or hearing impairments, their speech may be difficult to understand or hear due to physiological reasons.
    • Traditional voice conversion technologies typically require paired datasets of whispered and normal speech, which are costly to train and user-dependent.
  • Why This Problem is Important:

    • The low social acceptance of speech interaction technologies necessitates the development of more private, effective, and accessible solutions.
    • Improving the quality of whisper-to-normal voice conversion can enhance communication for individuals with hearing or speech impairments, increasing the societal value of the technology.
  • Motivation and Related Work:

    • Existing whisper speech recognition technologies face challenges such as high device requirements, difficulty in building datasets, and limited vocabulary recognition.
    • Specialized technologies (e.g., SilentVoice) have introduced innovative interaction methods but still require specialized hardware and training data.
    • The authors propose a novel solution to address these technical barriers in traditional whisper-to-normal voice conversion.

Solution

  • Proposed Method or Solution: The authors propose a zero-shot, real-time whisper-to-normal voice conversion mechanism called WESPER, based on a self-supervised learning framework. It consists of two main modules:

    • STU Encoder: Converts whispered or normal speech into shared speech units.
    • UTS Decoder: Reconstructs target speech (normal voice) from the shared speech units.
  • Innovative Features:

    • Does not require paired whispered and normal speech datasets; training can be performed using unpaired and unlabeled data.
    • The conversion process is speaker-independent, eliminating the need for user-specific training.
    • Maintains natural prosody and improves the quality of converted speech.
  • Implementation Steps and Key Techniques:

    1. STU Encoder Training:
      • Employs the HuBERT model for self-supervised pretraining to bridge the gap between whispered and normal speech.
      • Generates shared speech units using unpaired whispered and normal speech data.
      • Provides high-dimensional vectors to represent speech features.
    2. UTS Decoder Training:
      • Based on a modified FastSpeech2 model, eliminating the need for text annotation of speech.
      • Reconstructs target speech directly from shared speech units generated by the STU encoder.
      • Utilizes HiFi-GAN to produce high-quality audio.

Research Outcomes

  • Specific Achievements:

    • Introduced a real-time, zero-shot voice conversion technology and validated its effectiveness for normal speech and special speech impairment scenarios.
    • Compared to traditional methods, WESPER significantly reduces the gap between whispered and normal speech, with experiments demonstrating improved speech quality.
  • Advantages Over Existing Solutions:

    • Does not rely on paired whispered and normal speech datasets or text annotations.
    • Capable of performing voice conversion in real-time.
    • Demonstrated significant improvements for both normal users and individuals with hearing and speech impairments in experiments.
  • Experimental or Evaluation Results:

    • Whisper-to-Normal Voice Conversion:
      • Evaluated using Mean Opinion Score (MOS) and MUSHRA, confirming that WESPER significantly improves the quality of converted speech compared to whispered input, while maintaining natural prosody.
    • Speech Reconstruction and Tests for Individuals with Impairments:
      • For individuals with vocal cord polyps or spasmodic dysphonia, MOS and MUSHRA scores showed significant improvements in speech quality and naturalness.
      • Tests on speech reconstruction for individuals with hearing impairments indicated some quality improvements, though prosody consistency remained weak.
  • Limitations and Future Directions:

    • For the issue of inconsistent prosody in speech from individuals with hearing impairments, future work could explore small-scale fine-tuning for user customization.
    • Investigate more sophisticated speech input devices, such as non-audible low-frequency vibration detectors, to further optimize speech capture.
    • Examine its multilingual conversion capabilities, particularly for non-English speech datasets.
    • Explore integration with other AI or human-computer interaction systems to enhance user experience and technology usability.

Through this research, WESPER demonstrates the significant potential of artificial intelligence and speech interaction technologies in addressing physiological challenges and advancing privacy solutions in society.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/96115/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3580706
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
1 authors
sell
Subtopics
Intelligent Voice Assistants (Alexa, Siri, etc.), Voice Accessibility
work
Professions
Speech-Language Pathologists & Audiologists
article
Content Status
Full text indexed
hub
Related Papers
4 related papers