Creating Inclusive Voices for the 21st Century: A Non-Binary Text-to-Speech for Conversational Assistants

Multilingual & Cross-Cultural Voice InteractionAgent Personality & Anthropomorphism

Document Title

Creating Inclusive Voices for the 21st Century: A Non-Binary Text-to-Speech Technology for Conversational Assistant Development

Document Information

  • Subject Areas: Human-Computer Interaction, Speech Generation Technology, Non-Binary Gender Inclusivity
  • Keywords: Gender, Text-to-Speech (TTS), Voice Assistants, Voice User Interface, HCI Design, Inclusive Technology, LGBTQ+, Non-Binary Voices, Gender Attribution, Speech Synthesis

Research Background and Issues

  • Identified Problems or Challenges:

    1. Most current voice assistants (e.g., Siri and Alexa) default to female voices or offer only traditional male or female voice options. This binary gender design reinforces gender stereotypes and negative behaviors (e.g., sexual harassment or gender-based hostility).
    2. The lack of non-binary or gender-ambiguous voices results in insufficient representation of non-binary communities.
    3. Existing studies suggest that listeners perceive ambiguous gender voices as "untrustworthy or unfriendly," but these findings may no longer align with modern society's diverse gender culture.
  • Significance: As voice assistants become increasingly integrated into daily life, non-inclusive designs may harm the experiences and sense of belonging of marginalized groups. Moreover, inclusive technology design can promote gender equality and multicultural acceptance, driving societal progress through technological innovation.

  • Research Motivation and Related Work:

    1. Organizations like UNESCO have highlighted that gendered voice assistants may reinforce stereotypes of "female subservience."
    2. Previous attempts (e.g., the gender-neutral voice "Q") remained conceptual and lacked practical applications in speech synthesis.
    3. Representing non-binary and transgender voices requires new technical approaches to reflect the diversity of real-world expressions.

Proposed Solution

  • Proposed Solution:

    1. Developed a method for creating non-binary text-to-speech (TTS) voices, incorporating feedback from non-binary and transgender communities.
    2. Created a sample voice named "Sam" using this technology, with open-source sharing of the generation process and data to encourage further research.
    3. Conducted large-scale user surveys and comparative analyses on the acceptance and preferences for non-binary voice assistants among users of different gender identities.
  • Innovative Contributions:

    1. Went beyond simple pitch adjustments, expanding voice features to include intonation, rhythm, and vocal style across multiple dimensions.
    2. Emphasized collaboration with non-binary and transgender communities, integrating their needs and voice representation throughout the design and evaluation process.
    3. Open-sourced the TTS model, data, and processes, providing tools for others to develop customized non-binary voice options, enhancing technological sustainability and community participation.
  • Implementation Steps:

    1. Voice Actor Selection: Conducted an open call to identify suitable voice actors, with experts screening candidates based on technical and non-binary style criteria.
    2. Voice Data Production: Recorded hours of high-quality voice data using standard text, while expanding the sample library with publicly available audio from non-binary and transgender communities for secondary model training.
    3. Text-to-Speech Model Development: Utilized two TTS engines (proprietary Neural TTS and open-source Idlak) and applied transfer learning techniques to adjust trained feature models to the target non-binary range.
    4. Feedback and Iterative Design: Conducted two rounds of surveys within non-binary communities and a broader user base, collecting feedback and optimizing the generated voice through multiple iterations.
    5. Final Model Generation and Research: Finalized the voice profile and conducted large-scale user research to analyze emotional responses and applicability across different gendered voices.

Research Outcomes

  • Specific Outcomes:

    1. Proposed an integrated process combining speech synthesis methods and user research to generate and evaluate non-binary TTS voices.
    2. Designed a non-binary voice named "Sam," which can be directly applied as a voice option for conversational assistants.
    3. Conducted a large-scale survey with 1,010 participants, finding that non-binary voices were more appealing and trustworthy to non-binary individuals but had relatively lower acceptance among mainstream users.
  • Advantages Compared to Existing Solutions:

    1. Open-source and transparent: Unlike other closed projects, the developed data and models are publicly available, encouraging adoption and improvement by the developer community.
    2. Detailed focus: Iterative design based on community feedback ensures the generated voice closely aligns with the diverse expressions of non-binary identities.
    3. Multi-dimensional integration: Beyond pitch adjustments, the approach considers rhythm, semantic choices, and usage scenarios.
  • Experimental or Evaluation Results:

    1. Large-scale studies showed that non-binary voices enhanced reputation and applicability within specific user groups (e.g., non-binary and LGBTQ+ individuals).
    2. In certain scenarios, non-binary voices demonstrated strong impressions of intelligence and trustworthiness.
    3. While users across age groups and gender identities still showed a general preference for male voices, non-binary voices provided unique appeal for certain demographics.
  • Limitations and Future Directions:

    1. The study focused solely on American English, leaving other languages or dialects unaddressed in terms of cultural context and vocal habits.
    2. The open nature of the model may pose risks of misuse; ongoing attention to its societal impact and stricter restrictions on legitimate data use are necessary.
    3. Future work could expand to multilingual environments, incorporate diverse cultural contexts, and develop more varied voice features (e.g., multi-voice blending or richer emotional expressions).

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/96187/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581281
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Multilingual & Cross-Cultural Voice Interaction, Agent Personality & Anthropomorphism
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
7 related papers