Black LLMirror: User (Self) Perceptions in Black American English Interactions with LLMs
Authors
Paper Title
Black LLMirror: User (Self) Perceptions in Black American English Interactions with LLMs
Publication Info
- Topic area: User perceptions and self-perceptions in interactions with Black American English (BAE)-producing large language models (LLMs).
- Keywords: Black American English, dialect mirroring, large language models, user perception, self-esteem, code-switching, personalization, speech-based interaction, text-based interaction, linguistic bias.
Background and Problem
- Problem / challenge: Existing LLMs often fail to handle non-standard dialects like BAE effectively, leading to user dissatisfaction, inappropriate outputs, or perceived appropriation. Prior studies have focused on static or single-turn interactions and overlooked the impact on users' self-perception.
- Significance: Understanding how BAE speakers perceive and respond to LLMs using their dialect is critical for designing inclusive and effective language technologies, especially given the historical and cultural significance of BAE.
- Motivation and related work: Previous research has documented biases against BAE in NLP systems, user dissatisfaction with dialectal underperformance, and the psychological effects of interacting with biased systems. However, these studies have not explored multi-turn interactions or the effects on user self-perception, leaving a gap this paper addresses.
Solution
- Proposed approach: The study investigates BAE speakers' perceptions of BAE- and Standard American English (SAE)-producing LLMs across text- and voice-based interactions, focusing on both model perception and self-perception.
- Novelty:
- Examines multi-turn interactions with BAE- and SAE-producing LLMs.
- Measures both user self-perception and model perception across modalities.
- Incorporates qualitative analysis of open-ended user feedback.
- Validates BAE dialect density in LLM outputs using quantitative and perception-based methods.
- Procedure and key techniques:
- Participants (n=619) were recruited via Prolific, screened for BAE competence, and randomly assigned to interact with either a BAE- or SAE-producing LLM in text or voice.
- Pre- and post-interaction surveys measured self-perception (self-esteem, self-consciousness, affect) and model perception (warmth, competence, trust, sociability, comprehensibility).
- Open-ended feedback was analyzed thematically to capture nuanced user reactions.
- BAE dialect density in LLM outputs was validated using participant perceptions and an XGBoost-based dialect density predictor.
Results
- Concrete findings:
- Text-based BAE interactions significantly increased participants' self-esteem (d=0.116, p=0.014).
- Speech-based BAE interactions showed a slight, non-significant increase in negative feelings and self-consciousness.
- BAE models in speech were rated significantly lower than SAE models on warmth (p=0.012) and competence (p=0.01).
- Participants reported higher BAE usage when interacting with BAE models, corroborated by dialect density analysis (d=0.303, p=0.015).
- Advantage over baselines: The study highlights the nuanced effects of dialect mirroring, showing that BAE models can reduce the perceived need to code-switch but may evoke mixed or polarized reactions.
- Experiments / evaluation:
- Two-part surveys with randomized assignment to BAE or SAE models in text or voice.
- Statistical tests (paired t-tests, Mann-Whitney U-tests) and thematic analysis of open-text responses.
- Validation of BAE dialect density using participant perceptions and computational models.
- Limitations and future work:
- The study was underpowered due to sample size constraints, limiting subgroup analysis and domain-specific insights.
- Overrepresentation of certain demographic groups (e.g., women, participants aged 30–39) may have influenced findings.
- Future work could explore longitudinal interactions, other dialects, or more granular demographic effects.
Summary
This study investigates how native BAE speakers perceive themselves and BAE- or SAE-producing LLMs during multi-turn interactions in text and voice. Text-based BAE interactions significantly boosted self-esteem, while speech-based BAE interactions were rated lower on warmth and competence. Participants reported reduced code-switching with BAE models, but qualitative feedback revealed polarized reactions, ranging from affirming to offensive. These findings underscore the risks of default dialect mirroring and highlight the need for user-centered design principles, such as conservative personalization probing and explicit user feedback channels. The results contribute to understanding how LLMs can better serve linguistically marginalized communities.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 80%
“As an AI language model, I cannot”: Investigating LLM Denials of User Requests
CHI '24· Human-LLM Collaboration +1
- 80%
Responsibility Attribution in Human Interactions with Everyday AI Systems
CHI '25· AI Ethics, Fairness & Accountability +1
- 67%
Factors Influencing Perceived Fairness in Algorithmic Decision-Making: Algorithm Outcomes, Development Procedures, and Individual Differences
CHI '20· AI Ethics, Fairness & Accountability +1
- 67%
"Look! It's a Computer Program! It's an Algorithm! It's AI!'': Does Terminology Affect Human Perceptions and Evaluations of Algorithmic Decision-Making Systems
CHI '22· AI Ethics, Fairness & Accountability +2
- 67%
Envisioning and Understanding Orientations to Introspective AI: Exploring a Design Space with Meta.Aware
CHI '23· AI Ethics, Fairness & Accountability +2
- 67%
Generative AI and Perceptual Harms: Who’s Suspected of using LLMs?
CHI '25· Human-LLM Collaboration +2
- 60%
AI-Driven Mediation Strategies for Audience Depolarisation in Online Debates
CHI '24· Human-LLM Collaboration +2
- 60%
The Role of Inclusion, Control, and Ownership in Workplace AI-Mediated Communication
CHI '24· Human-LLM Collaboration +2
- 60%
Continual Human-in-the-Loop Optimization
CHI '25· Human-LLM Collaboration
Based on Jaccard similarity of research subtopics & professions (≥60%)