“Are Compliments Bad Now?”: Comparing LLMs and Human Interpretations of Gender Microaggressions in the Workplace
Authors
Paper Title
“Are Compliments Bad Now?”: Comparing LLMs and Human Interpretations of Gender Microaggressions in the Workplace
Publication Info
- Topic area: Interpretation of gender microaggressions in workplace scenarios by humans and large language models (LLMs).
- Keywords: Gender microaggressions, workplace discrimination, large language models, interpretive plurality, feminist HCI, situated knowledge, epistemic injustice, AI bias, interpretive sensitivity, ambiguity.
Background and Problem
- Problem / challenge: Existing automated systems for detecting microaggressions often fail to account for the ambiguity and context-dependence of such interactions. LLMs, while capable of high detection rates, may not align with human interpretations, particularly those shaped by lived experiences of discrimination.
- Significance: Microaggressions have significant personal and organizational consequences, including burnout and attrition. Effective detection systems could support equity and inclusion efforts but must preserve interpretive plurality to avoid marginalizing non-dominant perspectives.
- Motivation and related work: Prior research has treated microaggression detection as a text classification problem, often assuming objective ground truths. However, interpretations of microaggressions vary with social identity and lived experience. Recent studies show LLMs reflect dominant cultural norms, raising concerns about their suitability for socially sensitive tasks. This paper addresses the gap in understanding how LLMs and humans differ in interpreting gender microaggressions.
Solution
- Proposed approach: A mixed-methods study comparing human and LLM interpretations of workplace gender microaggressions, focusing on ratings and rationales for ten dialogue scenarios.
- Novelty:
- Empirical evidence of significant differences in how humans and LLMs interpret gender microaggressions.
- Identification of patterns of alignment and divergence between LLMs and human subgroups based on gender identity and lived experience.
- Introduction of the distinction between categorical sensitivity (LLMs) and situated sensitivity (humans).
- Design implications for preserving interpretive plurality in automated systems.
- Procedure and key techniques:
- Collected numerical ratings (1–5 Likert scale) and open-text rationales from 141 human participants (stratified by gender identity and lived experience) and 7 LLM models.
- Scenarios included 8 containing gender microaggressions and 2 neutral controls, adapted from real-world accounts.
- Analyzed ratings using non-parametric statistical tests and rationales using reflexive thematic analysis and descriptive coding.
Results
- Concrete findings:
- LLMs consistently assigned high ratings (4.7–5.0) with minimal variance, while humans showed greater interpretive diversity (3.04–4.73), influenced by gender identity and lived experience.
- Humans acknowledged ambiguity (7.87% of rationales) and contextual factors, while LLMs rarely did (0.18% ambiguity cues).
- LLM rationales centered on generalized harm (92.5% mechanism descriptions), whereas human rationales often grounded judgments in specific contextual anchors.
- Advantage over baselines: LLMs achieved high detection rates but lacked the contextual grounding and interpretive nuance demonstrated by humans, particularly those with lived experience of microaggressions.
- Experiments / evaluation:
- Human participants (n=135) stratified into six subgroups based on gender and lived experience.
- LLMs included Claude, GPT, and Gemini families, evaluated under identical conditions.
- Statistical tests (Kruskal–Wallis, Mann–Whitney) revealed significant differences in ratings and rationales across groups.
- Limitations and future work:
- Limited generalizability due to specific LLM versions, prompt design, and focus on gender-based microaggressions.
- Small sample size for certain subgroups (e.g., males with lived experience).
- Future work could explore intersectional microaggressions, cultural variations, and fine-tuning LLMs for contextual sensitivity.
Summary
This study highlights fundamental differences between human and LLM interpretations of workplace gender microaggressions. While LLMs exhibit categorical sensitivity by consistently identifying harm based on predefined categories, humans demonstrate situated sensitivity by grounding their judgments in contextual and relational factors. These findings challenge the assumption that higher detection rates equate to greater sensitivity and emphasize the importance of preserving interpretive plurality in automated systems. The paper proposes design implications for fostering reflective dialogue and avoiding algorithmic verdicts that override human judgment, advancing feminist HCI principles and critical computing methodologies.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
Funding AI for Good: A Call for Meaningful Engagement
CHI '26· AI Ethics, Fairness & Accountability +2
- 71%
Skin-Deep Bias: How Avatar Appearances Shape Perceptions of AI Hiring
CHI '26· AI Ethics, Fairness & Accountability +2
- 67%
Understanding the Boundaries between Policymaking and HCI
CHI '19· Algorithmic Fairness & Bias +1
- 63%
To Live in Their Utopia: Why Algorithmic Systems Create Absurd Outcomes
CHI '21· AI Ethics, Fairness & Accountability +2
- 63%
Regulating AI: Where U.S. State Policy and HCI (Mis)align
CHI '26· AI Ethics, Fairness & Accountability +3
- 63%
Mind in the Machine? Cross-Disciplinary Perceptions of Consciousness in Artificial Intelligence
CHI '26· Human-LLM Collaboration +3
- 63%
From the Field to the Algorithm: Understanding Indian Ethnographers' Perspectives on Responsible AI
CHI '26· AI Ethics, Fairness & Accountability +3
- 63%
Participatory AI Justice in HCI: A Scoping Review
CHI '26· Participatory Design +3
Based on Jaccard similarity of research subtopics & professions (≥60%)