“Are Compliments Bad Now?”: Comparing LLMs and Human Interpretations of Gender Microaggressions in the Workplace

AI Ethics, Fairness & AccountabilityAlgorithmic Fairness & BiasTechnology Ethics & Critical HCIUniversity Professors & ResearchersHCI ResearchersPrivacy Policy Makers

Paper Title

“Are Compliments Bad Now?”: Comparing LLMs and Human Interpretations of Gender Microaggressions in the Workplace

Publication Info

  • Topic area: Interpretation of gender microaggressions in workplace scenarios by humans and large language models (LLMs).
  • Keywords: Gender microaggressions, workplace discrimination, large language models, interpretive plurality, feminist HCI, situated knowledge, epistemic injustice, AI bias, interpretive sensitivity, ambiguity.

Background and Problem

  • Problem / challenge: Existing automated systems for detecting microaggressions often fail to account for the ambiguity and context-dependence of such interactions. LLMs, while capable of high detection rates, may not align with human interpretations, particularly those shaped by lived experiences of discrimination.
  • Significance: Microaggressions have significant personal and organizational consequences, including burnout and attrition. Effective detection systems could support equity and inclusion efforts but must preserve interpretive plurality to avoid marginalizing non-dominant perspectives.
  • Motivation and related work: Prior research has treated microaggression detection as a text classification problem, often assuming objective ground truths. However, interpretations of microaggressions vary with social identity and lived experience. Recent studies show LLMs reflect dominant cultural norms, raising concerns about their suitability for socially sensitive tasks. This paper addresses the gap in understanding how LLMs and humans differ in interpreting gender microaggressions.

Solution

  • Proposed approach: A mixed-methods study comparing human and LLM interpretations of workplace gender microaggressions, focusing on ratings and rationales for ten dialogue scenarios.
  • Novelty:
    1. Empirical evidence of significant differences in how humans and LLMs interpret gender microaggressions.
    2. Identification of patterns of alignment and divergence between LLMs and human subgroups based on gender identity and lived experience.
    3. Introduction of the distinction between categorical sensitivity (LLMs) and situated sensitivity (humans).
    4. Design implications for preserving interpretive plurality in automated systems.
  • Procedure and key techniques:
    • Collected numerical ratings (1–5 Likert scale) and open-text rationales from 141 human participants (stratified by gender identity and lived experience) and 7 LLM models.
    • Scenarios included 8 containing gender microaggressions and 2 neutral controls, adapted from real-world accounts.
    • Analyzed ratings using non-parametric statistical tests and rationales using reflexive thematic analysis and descriptive coding.

Results

  • Concrete findings:
    • LLMs consistently assigned high ratings (4.7–5.0) with minimal variance, while humans showed greater interpretive diversity (3.04–4.73), influenced by gender identity and lived experience.
    • Humans acknowledged ambiguity (7.87% of rationales) and contextual factors, while LLMs rarely did (0.18% ambiguity cues).
    • LLM rationales centered on generalized harm (92.5% mechanism descriptions), whereas human rationales often grounded judgments in specific contextual anchors.
  • Advantage over baselines: LLMs achieved high detection rates but lacked the contextual grounding and interpretive nuance demonstrated by humans, particularly those with lived experience of microaggressions.
  • Experiments / evaluation:
    • Human participants (n=135) stratified into six subgroups based on gender and lived experience.
    • LLMs included Claude, GPT, and Gemini families, evaluated under identical conditions.
    • Statistical tests (Kruskal–Wallis, Mann–Whitney) revealed significant differences in ratings and rationales across groups.
  • Limitations and future work:
    • Limited generalizability due to specific LLM versions, prompt design, and focus on gender-based microaggressions.
    • Small sample size for certain subgroups (e.g., males with lived experience).
    • Future work could explore intersectional microaggressions, cultural variations, and fine-tuning LLMs for contextual sensitivity.

Summary

This study highlights fundamental differences between human and LLM interpretations of workplace gender microaggressions. While LLMs exhibit categorical sensitivity by consistently identifying harm based on predefined categories, humans demonstrate situated sensitivity by grounding their judgments in contextual and relational factors. These findings challenge the assumption that higher detection rates equate to greater sensitivity and emphasize the importance of preserving interpretive plurality in automated systems. The paper proposes design implications for fostering reflective dialogue and avoiding algorithmic verdicts that override human judgment, advancing feminist HCI principles and critical computing methodologies.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222320/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790436
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
AI Ethics, Fairness & Accountability, Algorithmic Fairness & Bias, Technology Ethics & Critical HCI
work
Professions
University Professors & Researchers, HCI Researchers, Privacy Policy Makers
article
Content Status
Full text indexed
hub
Related Papers
8 related papers