CLARIS: Clear and Intelligible Speech from Whispered and Dysarthric Voices

Voice AccessibilityVibrotactile Feedback & Skin StimulationVoice User Interface (VUI) DesignSpeech-Language Pathologists & AudiologistsPsychiatrists & PsychotherapistsPhysicians, Nurses & Clinicians

Paper Title

CLARIS: Clear and Intelligible Speech from Whispered and Dysarthric Voices

Publication Info

  • Topic area: Speech-to-speech restoration for whispered and dysarthric voices
  • Keywords: Whispered speech, dysarthric speech, speech restoration, voice conversion, personalization, multilingual, prosody, intelligibility, autoregressive modeling, data augmentation

Background and Problem

  • Problem / challenge: Existing speech-to-speech systems struggle with whispered and dysarthric speech due to limitations in generalization, reliance on hand-crafted data augmentation, and inability to handle severe speech impairments or cross-lingual scenarios. Personalization and alignment-free modeling remain underexplored.
  • Significance: Whispered and dysarthric speech hinder effective communication and accessibility, creating barriers for individuals in both private and professional contexts. Addressing these challenges can enable inclusive and natural voice interactions.
  • Motivation and related work: Prior work, such as WESPER and DistillW2N, has made progress in whisper-to-speech conversion using self-supervised and generative models. However, these systems often fail on unseen speakers, accents, and severe disorders. Dysarthric speech systems face high error rates in open-vocabulary settings, and existing augmentation methods produce unnatural artifacts. This paper builds on advances in self-supervised learning and autoregressive modeling to address these gaps.

Solution

  • Proposed approach: CLARIS (Clear and Accessible Restoration of Impaired Speech), an autoregressive speech-to-speech restoration framework that converts whispered and dysarthric input into natural, intelligible speech.
  • Novelty:
    1. Unified, alignment-free architecture for whispered and dysarthric speech conversion using an autoregressive transformer.
    2. Cross-lingual and clinically relevant personalization with minimal data (15–30 minutes).
    3. Integrated TTS-based augmentation and a Real–Synthetic Alignment Discriminator (RSAD) for robust training on synthetic data.
    4. Lightweight design (40.71M parameters) with real-time inference capabilities.
  • Procedure and key techniques:
    • Data augmentation: Synthetic atypical speech generated using VITS TTS, scaled from 15–30 minutes of real data to hundreds of hours.
    • AS2UT encoder: Extracts mel-spectrogram features and predicts linguistic units with auxiliary character decoders and CTC supervision.
    • RSAD: Aligns real and synthetic embeddings via adversarial training to mitigate domain mismatch.
    • Unit-to-speech renderer: Converts predicted units into natural speech using a HiFi-GAN-based vocoder.

Results

  • Concrete findings:
    • Achieved 12% WER on wTIMIT whispers and 31% WER on TORGO dysarthric speech, outperforming baselines like WESPER and DistillW2N.
    • Delivered 29.21% WER on Hindi whispers, demonstrating cross-lingual generalization.
    • Personalized models reduced WER to 12.63% for Indian-accent whispers and 38.21% for dysarthric speech with 30 minutes of fine-tuning data.
  • Advantage over baselines:
    • Outperformed ASR-TTS pipelines and state-of-the-art systems like WESPER and DistillW2N in intelligibility, BLEU, and ROUGE-L scores.
    • Significantly improved subjective ratings for quality, intelligibility, and prosody.
  • Experiments / evaluation:
    • Benchmarked on English whispers (wTIMIT), Hindi whispers, and dysarthric speech (TORGO).
    • Metrics included WER, CER, BLEU, ROUGE-L, and subjective MOS ratings.
    • Evaluated generalization to unseen speakers, accents, and languages, as well as personalization with limited data.
  • Limitations and future work:
    • Limited subjective evaluations with non-impaired listeners; future studies should involve individuals with speech impairments.
    • Autoregressive design limits inference speed compared to non-autoregressive models.
    • Focused on English and Hindi; future work should extend to other languages and disorders.

Summary

CLARIS is a lightweight, alignment-free speech restoration framework that converts whispered and dysarthric speech into intelligible, natural-sounding output. It achieves state-of-the-art performance across multiple datasets and languages, with strong generalization to unseen speakers and accents. Personalization with minimal data enables speaker-specific adaptation, making it suitable for inclusive and accessible voice interaction. Future work will focus on expanding to more languages, disorders, and real-world deployment scenarios.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222969/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791734
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Voice Accessibility, Vibrotactile Feedback & Skin Stimulation, Voice User Interface (VUI) Design
work
Professions
Speech-Language Pathologists & Audiologists, Psychiatrists & Psychotherapists, Physicians, Nurses & Clinicians
article
Content Status
Full text indexed
hub
Related Papers
1 related papers