Sonora: Human-AI Co-Creation of 3D Audio Worlds and its Impact on Anxiety and Cognitive Load

Generative AI (Text, Image, Music, Video)Smart Home Interaction DesignPsychiatrists & PsychotherapistsAthletes & Fitness EnthusiastsFreelancers (Design, Writing, Translation)

Research Background and Problem

  • What problems or challenges did the authors identify?
    Soundscapes are widely used for relaxation and emotional regulation, but their potential for interactivity and personalized experiences has not been fully explored. Additionally, traditional relaxation techniques often rely on "one-size-fits-all" fixed solutions or overly complex options, which may not be suitable for specific users, especially those with high anxiety.

  • Why is this problem important?
    Anxiety is one of the most common mental health disorders, affecting over 300 million people worldwide. Personalized mental health interventions have been shown to improve user engagement and effectiveness. Exploring the creation of dynamic, customizable sound environments can help provide more effective, non-invasive therapeutic methods.

  • Research Motivation and Related Work
    Inspired by existing sound therapy approaches and leveraging generative AI technologies (e.g., large language models and diffusion models), the authors propose a solution to enable real-time generation and personalization of soundscapes, enhanced through user interaction. While some studies have explored the benefits of sound on mental health, most focus on passive listening rather than interactive experiences. Additionally, research on generative virtual worlds has predominantly concentrated on the visual domain, and this study extends it to the auditory dimension.

Solution

  • What methods or solutions did the authors propose?
    The authors developed an AI-driven system called "Sonora," which allows users to create and navigate 3D auditory worlds through voice commands. The system integrates large-scale language models (LLMs), audio diffusion models, and the Unity3D game engine to support real-time sound generation, spatialization, and on-demand personalization.

  • What are the innovative aspects of this solution?

    • No Graphical User Interface: The system provides a fully voice-based interaction experience, avoiding visual overload and making it suitable for specific contexts such as sleep aids or use during driving.
    • Real-time Generation and Editing: Users can dynamically add, remove, or adjust sound elements, ensuring a seamless and personalized experience.
    • Innovative System Architecture: The system integrates task-specific LLM modules for sound selection, semantic grouping, and 3D sound source positioning, ensuring that each sound aligns with the user's expected real-world scenarios.
  • What are the implementation steps and key technologies used?

    • Users input voice commands, which are corrected and interpreted by the Prompt Interpreter module.
    • The Planner module determines which LLM module to invoke (e.g., creating a new scene or removing a specific sound).
    • The SoundGroupCreator or SoundSingleCreator modules generate and spatialize sounds.
    • Dynamic animations (e.g., birds circling in the sky) are introduced to enhance immersion.
    • AI voice guidance is provided via the Azure Speech-to-Text module to ensure smooth user interaction.
      The system utilizes GPT-4o (LLM) and StabilityAI audio diffusion models to generate unique, high-quality sounds.

Research Outcomes

  • What specific results were achieved?

    • Developed and validated the Sonora system, marking the first use of generative AI to construct interactive, customizable 3D soundscapes.
    • Experiments demonstrated that Sonora's interactivity and personalization features significantly enhanced enjoyment without increasing users' cognitive load.
    • Highlighted the value of personalized experiences for mental health interventions, particularly in attracting individuals with anxiety.
  • What advantages does it have compared to existing solutions?

    • Provides genuine interactivity: Users can create unique sound worlds rather than passively listening to pre-set materials.
    • Simplifies interaction: The voice-driven interface eliminates the need for complex graphical operations, reducing cognitive barriers.
    • Real-time functionality: Sound generation and editing enable smooth transitions in dynamic scenes.
  • What were the experimental or evaluation results?

    • Participants: 32 individuals (including high-anxiety and low-anxiety groups).
    • Anxiety Changes: Anxiety levels significantly decreased in both the Sonora users and the control group, particularly among high-anxiety users.
    • User Experience: Sonora scored significantly higher in enjoyment compared to passive soundscapes, with users highly rating its personalization features.
    • Cognitive Load and Physiological Measurements: No significant differences were found between the groups in terms of cognitive load, heart rate, or heart rate variability, indicating that Sonora's interactive features did not increase user burden.
  • Limitations and Future Directions

    • Sound Quality: The realism of certain sounds generated by diffusion models (e.g., music and human voices) is relatively low, and improving generation quality is a key future goal.
    • Interaction Optimization: Enhancing speech recognition accuracy, especially for complex commands, is necessary.
    • Sample Size: The current sample size is small; future studies should expand the testing scope, particularly for long-term use by high-anxiety individuals.
    • Further Applications: Explore Sonora's potential in education, gaming, meditation, and its applicability across cultural contexts.

Conclusion

Sonora is the first system to combine LLMs and audio diffusion models for interactive 3D sound environments, significantly advancing the technology and application prospects in this field. Its personalized features and intuitive voice interaction demonstrate potential for anxiety relief and pave the way for future applications in health, education, and entertainment.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189494/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713316
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Smart Home Interaction Design
work
Professions
Psychiatrists & Psychotherapists, Athletes & Fitness Enthusiasts, Freelancers (Design, Writing, Translation)
article
Content Status
Full text indexed
hub
Related Papers
0 related papers