Inter(sectional) Alia(s): Ambiguity in Voice Agent Identity via Intersectional Japanese Self-Referents
Authors
Research Background and Issues
- Identified Issues: The authors observed ethical concerns arising from AI conversational agents mimicking human communication, particularly in how language and voice express social identities, which may lead to biases or stereotypes. Additionally, many studies have not thoroughly explored AI identity through the lens of multiple social identities, such as intersectional identity markers.
- Importance of the Issue: AI may inadvertently exacerbate biases related to gender, age, and social status, affecting users' trust and satisfaction with the agents. Such design biases could negatively impact human-computer interaction, reinforcing stereotypes or gender inequality.
- Research Motivation and Related Work: Existing research primarily focuses on algorithmic bias but lacks sufficient investigation into intersectional perspectives of AI identity perception. In Japanese, self-referential expressions (e.g., pronouns) not only involve gender but also carry significant social identity implications, including age, formality, and regional characteristics.
Solution
- Research Methodology: The authors designed a crowdsourced experiment using three ChatGPT voices (Juniper, Breeze, and Ember) and seven Japanese self-referential expressions (including gender-neutral and intersectional options) to explore how combinations of these voices and expressions influence users' perceptions of social identity.
- Innovations:
- For the first time, combining voice with Japanese intersectional self-referential expressions to study AI identity perception.
- Introducing non-pronoun self-references (e.g., using names instead of pronouns) as a daily and exploratory form of self-expression.
- Investigating specific combinations that help evade or blur gender perception.
- Implementation Steps:
- Generating short videos with different voice and self-referential expression combinations using ChatGPT.
- Fixing sentence-ending expressions in each video to avoid introducing additional identity signals through speech style.
- Collecting evaluation data from 204 native Japanese speakers, including perceptions of age, gender, formality, regionality, and "cuteness."
Research Findings
- Specific Findings:
- Gender perception was strongly influenced by voice genderization; Breeze (expected to be gender-neutral) was perceived as masculine by most participants.
- Intersectional self-referential expressions (e.g., "ぼく," "あたし") successfully blurred traditional binary gender perceptions.
- Certain combinations (e.g., Juniper using "ぼく," Ember using "あたし") elicited impressions of gender ambiguity or even "gender crossing."
- Advantages:
- Providing direct evidence of the combined effects of language and voice on social identity perception, showcasing a culturally sensitive and actionable approach to AI identity design.
- Highlighting the importance of understanding identity perception through intersectional theory compared to existing studies, enabling the design of neutral or diverse AI personalities.
- Experimental or Evaluation Results:
- Gender ambiguity was more pronounced when using gendered voices paired with non-traditional self-referential expressions, such as Juniper with "ぼく."
- The "ぼくっこ" impression of Juniper was significantly associated with "cuteness" ratings, indicating that the combination of voice and language can trigger specific cultural and emotional perceptions.
- Limitations and Future Directions:
- Breeze's voice failed to achieve gender neutrality, highlighting the need for developing truly gender-neutral voices in the future.
- Regional impressions were weak in the experiment, suggesting future studies should incorporate dialects or tonal variations.
- Interactive studies are needed to validate more dynamic changes in identity perception during actual user interactions.
Conclusion and Recommendations
This study provides pioneering research on the combination of Japanese intersectional self-referential expressions and AI voices, offering new theoretical foundations and design directions for voice user interfaces and AI design. Future work should focus on addressing gender-neutral voice design issues and conducting diversity analyses across more cultural and linguistic contexts to further enhance AI agents' ability to express social identities.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How do voice and Japanese first-person pronoun combinations in AI conversational agents affect users' perception of social identity?Category: Gender, Sexuality Bias, and Women/LGBTQ+ Experiences in AI, Technology, and Online PlatformsSimilar questionsarrow_forward
- Which pronoun and voice pairings can blur or cross binary gender perception boundaries?Category: Gender, Sexuality Bias, and Women/LGBTQ+ Experiences in AI, Technology, and Online PlatformsSimilar questionsarrow_forward
- How can cultural sensitivity be introduced in AI design to avoid reinforcing gender stereotypes?Category: Gender, Sexuality Bias, and Women/LGBTQ+ Experiences in AI, Technology, and Online PlatformsSimilar questionsarrow_forward
Practical Problems
1- Users' trust in AI agents may decrease due to voice or language bias.Category: Gender, Sexuality Bias, and Women/LGBTQ+ Experiences in AI, Technology, and Online PlatformsSimilar questionsarrow_forward
- 100%
The Sound of Support: Gendered Voice Agent as Support to Minority Teammates in Gender-Imbalanced Team
CHI '24· Intelligent Voice Assistants (Alexa, Siri, etc.) +2
- 67%
Emotional Dialogue Generation Using Image-Grounded Language Models
CHI '18· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 67%
Beyond Translation: Design and Evaluation of an Emotional and Contextual Knowledge Interface for Foreign Language Social Media Posts
CHI '18· Multilingual & Cross-Cultural Voice Interaction +1
- 67%
Understanding Affective Experiences with Conversational Agents
CHI '19· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 67%
A New Uncanny Valley? The Effects of Speech Fidelity and Human Listener Gender on Social Perceptions of a Virtual-Human Speaker
CHI '22· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 67%
Simulating Human Imprecision in Temporal Statements of Intelligent Virtual Agents
CHI '22· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 67%
Creating Inclusive Voices for the 21st Century: A Non-Binary Text-to-Speech for Conversational Assistants
CHI '23· Multilingual & Cross-Cultural Voice Interaction +1
- 67%
Super Kawaii Vocalics: Amplifying the “Cute” Factor in Computer Voice
CHI '25· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 67%
Exploring Humor as a Repair Strategy During Communication Breakdowns with Voice Assistants
CUI '23· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 67%
“Hello, This is a Voice Assistant Calling" When a Human Voice Calls Claiming to Be a Machine on an Ordinary Day
DIS '25· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)