"A Voice that Suits the Situation": Understanding the Needs and Challenges for Supporting End-User Voice Customization

Agent Personality & AnthropomorphismGenerative AI (Text, Image, Music, Video)

Document Title

“A Voice that Suits the Situation”: Understanding the Needs and Challenges for Supporting End-User Voice Customization

Document Information

  • Subject Area: Human-Computer Interaction and Voice Customization Technology
  • Keywords: Voice Perception, Voice Customization, User Experience, Voice Quality, Anonymous Interaction, Virtual Worlds

Research Background and Issues

  • What problems or challenges did the authors identify?

    • The demand for voice customization in current virtual environments remains under-addressed, with existing technologies focusing more on appearance customization while neglecting the impact of voice on user evaluation and interaction.
    • There is a gap in data and design frameworks regarding how existing voice customization tools are used and what users need from them.
  • Why is this issue important?

    • In an era of increased social distancing and prevalent virtual interactions, voice and appearance together constitute a user’s identity in virtual worlds. Voice customization can enhance user experience and improve self-representation.
    • Voice customization has potential value in protecting user identity privacy and improving the quality of verbal expression, particularly in online gaming and educational contexts.
  • Research Motivation and Related Work:

    • Research in the metaverse domain has revealed a widespread demand for customizable virtual appearances, and related literature indicates that voice significantly influences evaluations and impressions in interpersonal interactions.
    • Voice customization enables personalized expression while protecting privacy, but existing technologies primarily focus on simple filter-based modifications and fail to provide comprehensive voice-shaping capabilities.

Solutions

  • What methods or solutions did the authors propose?

    • Conducted online surveys and semi-structured interviews to study users’ voice preferences and customization needs in various scenarios.
    • Developed a web-based voice customization prototype that allows users to search for and select target voices.
  • What is innovative about this solution?

    • Examined users’ diverse needs for voice expression in specific scenarios.
    • Provided multi-dimensional voice attribute selection features, including breathiness, smoothness, hoarseness, and variability, to help users find suitable voices.
  • What were the implementation steps? What key technologies were used?

    1. Online Survey Study: Investigated voice customization needs and satisfaction, collecting user preferences for voice customization in different scenarios.
    2. Prototype Development and Interviews:
      • Developed a web application prototype based on the Django framework for voice searching and filtering.
      • Used the Speech Accent Archive dataset and employed Amazon Mechanical Turk to annotate voice attributes.
      • Conducted interviews to gather user feedback on selection criteria and the voice customization system.

Research Findings

  • What specific results were achieved?

    • Users demonstrated a strong demand for voice customization, especially in online communication scenarios.
    • Users preferred different voice styles depending on the interaction context, such as prioritizing anonymity or enhancing expression quality.
    • Proposed design recommendations for voice customization tools, including support for multiple voice attributes, multi-level options, and context-based profiles.
  • How does it compare to existing solutions?

    • The study revealed detailed needs for voice customization in specific scenarios, such as hiding gender in gaming or seeking professional voice expression in work settings.
    • The prototype offered richer functionality than traditional filtering tools, allowing users to listen to and select specific voice samples.
  • What were the experimental or evaluation results?

    • Over 80% of survey participants expressed a desire to change their voice in at least one scenario.
    • Scenarios requiring anonymous interaction (e.g., online gaming and customer service) showed the highest demand for voice customization.
    • Semi-structured interviews further confirmed users’ specific attribute needs for voice customization in different scenarios, such as modular voice changes to maintain speech recognizability.
  • Limitations and Future Directions

    • Limitations:

      • Data primarily came from Korean users, which may not generalize to other cultural contexts.
      • The prototype system had limited diversity in voice samples, and voice annotations were subjective.
      • Real-time voice transformation functionality was not implemented.
    • Future Directions:

      • Develop real-time voice synthesis capabilities to meet users’ needs for live interactions.
      • Expand the voice dataset and make it publicly available to support broader languages and cultures.
      • Explore the design of voice customization technologies for user privacy protection and crime prevention (e.g., voice forgery).

Design Recommendations

  • Support adjustments for various voice attributes, such as smoothness and variability, while introducing additional parameters like pitch, speed, and volume.
  • Add emotional or personalized tag-based search features, such as “#professional” or “#happy.”
  • Provide optimized voice suggestions similar to the user’s original voice to reduce interaction awkwardness and enhance recognizability.
  • Support the creation of scenario-based voice profiles, allowing users to quickly switch to suitable voice settings.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/68931/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/abs/10.1145/3491102.3501856
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Agent Personality & Anthropomorphism, Generative AI (Text, Image, Music, Video)
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
3 related papers