Leveraging AI-Generated Emotional Self-Voice to Nudge People towards their Ideal Selves

Intelligent Voice Assistants (Alexa, Siri, etc.)Generative AI (Text, Image, Music, Video)AI Ethics, Fairness & AccountabilityPsychiatrists & Psychotherapists

Research Background and Issues

  • Identified Problems or Challenges:
    Emotions significantly influence decision-making and goal pursuit, yet many individuals face difficulties in visualizing their ideal selves and adopting traditional cognitive-behavioral intervention methods. Moreover, the reliance on mental imagery in traditional cognitive-behavioral techniques may not provide adequate support for individuals who lack visual imagination.

  • Significance of the Research:
    The concept of the ideal self is a critical driver of intrinsic motivation, helping individuals adjust their behavior to bridge the gap between their "current self" and "ideal self." Since perceptions of the ideal self can foster long-term behavioral change, developing methods to effectively evoke such perceptions is of great importance.

  • Research Motivation and Related Work:
    Recent advancements in generative models, such as voice generation and cloning technologies, have made it possible to create more realistic self-representations. Existing studies have shown that the self-similarity of virtual avatars' appearance and voice can influence individuals' behavior and cognition. However, behavior interventions based on self-voice remain in the early stages of research and have not explored methods combining personalized emotional voice with ideal self systems.


Solution

  • Proposed Method or Solution:
    The authors propose the Emotional Self-Voice (ESV) system, which integrates emotional language models (LLM) and voice cloning technology to generate ideal self-responses using the individual's own voice. These responses aim to encourage individuals to align with their ideal selves through emotional expression and personalized characteristics.

  • Innovative Aspects of the Solution:

    1. Introduced a novel approach combining emotional LLM with voice cloning technology.
    2. Dynamically generates emotional, self-voice-specific audio outputs rather than relying on static voice recordings.
    3. The system emphasizes not only content generation but also the emotional quality of the voice, enhancing the intervention's immersion and engagement.
  • Implementation Steps and Key Technologies:

    1. Collect user-provided scenarios and descriptions of their ideal self.
    2. Use GPT-4 to generate customized text, designing natural language outputs with emotional expressions.
    3. Employ Hume AI tools to synthesize emotionalized speech.
    4. Utilize ElevenLabs voice cloning technology to convert the generated speech into personalized audio resembling the user's voice.
    5. Provide users with options to adjust the emotional characteristics of the generated output (e.g., positivity and intensity of emotional expression).

Research Outcomes

  • Specific Results:

    1. Overall Positive Effects: All three conditions (static imagination, text generation, ESV) effectively increased emotional positivity, resilience, confidence, and motivation.
    2. Uniqueness of ESV: ESV was rated as a unique, personalized, and highly engaging intervention method.
    3. Ideal Self Representation: Voice generation was perceived as a more vivid way to present the ideal self, helping users reduce the cognitive gap between their current and ideal selves.
  • Advantages Compared to Existing Solutions:

    1. Provides dynamic, emotional voice feedback instead of static text or pre-recorded audio.
    2. Enhances user experience through personalization and relatability (via voice cloning technology).
    3. Adapts to varying scenarios, such as handling goal failure and habit formation.
  • Experimental or Evaluation Results:

    1. Users in the ESV group reported higher familiarity and naturalness with the generated voice, along with more significant emotional effects.
    2. In future-oriented scenarios (e.g., habit formation), ESV demonstrated a notable positive impact on motivation.
    3. Emotional analysis showed that the generated voice content achieved a good balance between positivity and authenticity.
  • Limitations and Future Directions:

    1. Limited to Short-Term Effects: The study only examined immediate emotional and cognitive outcomes, without tracking long-term behavioral changes.
    2. Potential Habituation Risk: Prolonged exposure may lead to user desensitization to intervention content or voice.
    3. Expansion Possibilities: The authors suggest exploring multimodal interventions (e.g., combining images or virtual reality) and non-self voices (e.g., voices of ideal others) in future research.
    4. Privacy and Ethical Concerns: The use of voice cloning technology necessitates safeguards against misuse for deception or other malicious purposes.

Conclusion

This study introduces an emotional intervention tool (the ESV system) that combines generative models and voice cloning technology to construct and deliver ideal self-voices. The system significantly improves individuals' emotional and cognitive performance in two critical scenarios: goal failure and habit formation. Future work could further refine this innovative system by tracking long-term intervention effects, expanding multi-layered application scenarios, and optimizing privacy protection mechanisms.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189268/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713359
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Intelligent Voice Assistants (Alexa, Siri, etc.), Generative AI (Text, Image, Music, Video), AI Ethics, Fairness & Accountability
work
Professions
Psychiatrists & Psychotherapists
article
Content Status
Full text indexed
hub
Related Papers
0 related papers