The Sound of Hallucinations: Toward a more convincing emulation of internalized voices

Honorable Mention
Human Pose & Activity RecognitionConversational ChatbotsAgent Personality & AnthropomorphismPsychiatrists & PsychotherapistsHCI Researchers

Title of the Paper

The Sound of Hallucinations: Toward a more convincing emulation of internalized voices

Bibliographic Information

  • Research Domain: Human-Computer Interaction, Speech Processing, Development of Psychological Therapy Tools
  • Keywords: Avatar Therapy, Voice Transformation Interface, Speech Synthesis, Voice Space Navigation, Mixed Reality

Research Background and Problem Statement

  • Problem or Challenge:

    • Schizophrenia patients often experience internalized voices (hallucinations), which lack external references and are difficult to replicate.
    • Existing voice manipulation tools are complex and limited in capability, only allowing minor adjustments and failing to realistically reproduce the voices patients hear in their minds.
    • As augmented reality technology is increasingly applied in therapeutic contexts, creating high-quality virtual voices remains a challenge.
  • Significance:

    • Treatments for schizophrenia, such as Avatar Therapy, can alleviate symptoms by enabling patients to converse with virtual avatars. However, if the avatar's voice fails to convincingly simulate the patient's hallucinations, the therapeutic effect may be compromised.
  • Motivation:

    • To develop an easy-to-use yet powerful voice generation framework that does not require patients or therapists to have expertise in voice manipulation.
    • To provide greater expressiveness in voice control to address the limitations of current tools.

Solution

  • Proposed Method or Solution:

    • A voice generation framework was designed, utilizing Principal Component Analysis (PCA) to identify the primary dimensions of variation in voice feature vectors, which serve as control parameters.
    • A user-friendly interface was provided, allowing exploration of existing voice samples and further voice modifications.
    • Two voice manipulation techniques were proposed:
      1. Voice Space Navigation: Identifying samples within a specific multidimensional format that more closely resemble the target voice imagined by the user.
      2. Voice Parameter Editing: Adjusting parameters such as pitch, resonance, hoarseness, and emotional prosody to match the target voice.
      3. Voice Blending Technique: Generating new voices from two selected samples through linear interpolation, expanding the range of selectable voice samples.
  • Innovations:

    • Introducing visualization and navigation in a low-dimensional voice feature space to help users intuitively locate their target voice.
    • Providing direct modification methods for auditory characteristics (e.g., hoarseness, resonance) without requiring complex knowledge of audio domains.
    • Offering a simple and intuitive approach to voice blending while maintaining natural voice quality.
  • Implementation Steps and Techniques:

    • Extracting voice samples from the LibriSpeech corpus.
    • Using multi-voice models to extract voice feature vectors and projecting high-dimensional feature spaces into two dimensions via UMAP.
    • The interface allows users to search for initial voices, then precisely adjust relevant parameters or use interpolation to generate new voices.

Research Findings

  • Specific Results:

    • User experiments validated the effectiveness of the two voice generation techniques; compared to existing commercial voice transformation tools, the proposed methods more convincingly reproduced target voices.
    • Users preferred the proposed voice blending technique, citing its simplicity and ease of use.
  • Advantages:

    • Compared to traditional voice transformation tools, the generated voices were more natural, reducing issues of audio distortion.
    • Provided a solution for generating target voices without reference recordings, particularly useful for hallucination simulation scenarios.
  • Experimental or Evaluation Results:

    • Both techniques significantly improved the proximity of generated voices to target voices.
    • Through Multidimensional Scaling (MDS), users categorized generated voices closer to the target voice.
  • Limitations and Future Directions:

    • Limitations:

      • Voice editing parameters may have interdependent effects, where adjusting one parameter could simultaneously affect other characteristics.
      • The system currently focuses on North American English voices, with limited applicability to other languages or accents.
      • Computation time is relatively long, and real-time feedback for button adjustments has not yet been implemented.
    • Future Directions:

      • Enhance the balance of age and gender in the corpus to better match diverse voices.
      • Introduce language embedding techniques to support multilingual voice generation.
      • Optimize pre-generated voice samples to reduce interaction time.

Conclusion

This paper presents a voice generation framework for psychiatric therapy (Avatar Therapy). By exploring voice samples and adjusting parameters, the system can generate voices closer to those imagined by patients. This has significant potential for treating hallucination patients and advancing virtual avatar technologies, while addressing shortcomings in existing speech synthesis tools.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/68812/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/abs/10.1145/3491102.3501871
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2022
emoji_events
Award
Honorable Mention
group
Authors
5 authors
sell
Subtopics
Human Pose & Activity Recognition, Conversational Chatbots, Agent Personality & Anthropomorphism
work
Professions
Psychiatrists & Psychotherapists, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers