Inter(sectional) Alia(s): Ambiguity in Voice Agent Identity via Intersectional Japanese Self-Referents

Intelligent Voice Assistants (Alexa, Siri, etc.)Multilingual & Cross-Cultural Voice InteractionAgent Personality & Anthropomorphism

Research Background and Issues

  • Identified Issues: The authors observed ethical concerns arising from AI conversational agents mimicking human communication, particularly in how language and voice express social identities, which may lead to biases or stereotypes. Additionally, many studies have not thoroughly explored AI identity through the lens of multiple social identities, such as intersectional identity markers.
  • Importance of the Issue: AI may inadvertently exacerbate biases related to gender, age, and social status, affecting users' trust and satisfaction with the agents. Such design biases could negatively impact human-computer interaction, reinforcing stereotypes or gender inequality.
  • Research Motivation and Related Work: Existing research primarily focuses on algorithmic bias but lacks sufficient investigation into intersectional perspectives of AI identity perception. In Japanese, self-referential expressions (e.g., pronouns) not only involve gender but also carry significant social identity implications, including age, formality, and regional characteristics.

Solution

  • Research Methodology: The authors designed a crowdsourced experiment using three ChatGPT voices (Juniper, Breeze, and Ember) and seven Japanese self-referential expressions (including gender-neutral and intersectional options) to explore how combinations of these voices and expressions influence users' perceptions of social identity.
  • Innovations:
    1. For the first time, combining voice with Japanese intersectional self-referential expressions to study AI identity perception.
    2. Introducing non-pronoun self-references (e.g., using names instead of pronouns) as a daily and exploratory form of self-expression.
    3. Investigating specific combinations that help evade or blur gender perception.
  • Implementation Steps:
    1. Generating short videos with different voice and self-referential expression combinations using ChatGPT.
    2. Fixing sentence-ending expressions in each video to avoid introducing additional identity signals through speech style.
    3. Collecting evaluation data from 204 native Japanese speakers, including perceptions of age, gender, formality, regionality, and "cuteness."

Research Findings

  • Specific Findings:
    1. Gender perception was strongly influenced by voice genderization; Breeze (expected to be gender-neutral) was perceived as masculine by most participants.
    2. Intersectional self-referential expressions (e.g., "ぼく," "あたし") successfully blurred traditional binary gender perceptions.
    3. Certain combinations (e.g., Juniper using "ぼく," Ember using "あたし") elicited impressions of gender ambiguity or even "gender crossing."
  • Advantages:
    1. Providing direct evidence of the combined effects of language and voice on social identity perception, showcasing a culturally sensitive and actionable approach to AI identity design.
    2. Highlighting the importance of understanding identity perception through intersectional theory compared to existing studies, enabling the design of neutral or diverse AI personalities.
  • Experimental or Evaluation Results:
    1. Gender ambiguity was more pronounced when using gendered voices paired with non-traditional self-referential expressions, such as Juniper with "ぼく."
    2. The "ぼくっこ" impression of Juniper was significantly associated with "cuteness" ratings, indicating that the combination of voice and language can trigger specific cultural and emotional perceptions.
  • Limitations and Future Directions:
    1. Breeze's voice failed to achieve gender neutrality, highlighting the need for developing truly gender-neutral voices in the future.
    2. Regional impressions were weak in the experiment, suggesting future studies should incorporate dialects or tonal variations.
    3. Interactive studies are needed to validate more dynamic changes in identity perception during actual user interactions.

Conclusion and Recommendations

This study provides pioneering research on the combination of Japanese intersectional self-referential expressions and AI voices, offering new theoretical foundations and design directions for voice user interfaces and AI design. Future work should focus on addressing gender-neutral voice design issues and conducting diversity analyses across more cultural and linguistic contexts to further enhance AI agents' ability to express social identities.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188637/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713323
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Intelligent Voice Assistants (Alexa, Siri, etc.), Multilingual & Cross-Cultural Voice Interaction, Agent Personality & Anthropomorphism
work
Professions
article
Content Status
Full text indexed
hub
Related Papers
10 related papers