Writing with AI Can Reduce Gender Bias in Hiring Evaluations
Best PaperAuthors
Paper Title
Writing with AI Can Reduce Gender Bias in Hiring Evaluations
Publication Info
- Topic area: Gender bias reduction in hiring evaluations using AI writing assistants.
- Keywords: Gender bias, AI writing assistants, hiring evaluations, stereotypes, competence, warmth, decision-making, salary gap, language interventions.
Background and Problem
- Problem / challenge: Gender stereotypes persist in workplace evaluations, associating competence with men and warmth with women, leading to biased hiring and salary decisions. Existing interventions (e.g., awareness training, role models) often yield mixed results and rely on indirect mechanisms.
- Significance: Addressing gender bias in hiring is critical for workplace equity, as biases affect leadership opportunities, salary offers, and perceptions of competence.
- Motivation and related work: Previous research highlights the role of language in perpetuating stereotypes. AI writing assistants, which influence language production, present an opportunity to intervene directly by reshaping descriptive language. However, their potential for reducing bias has not been systematically studied.
Solution
- Proposed approach: An AI writing assistant that provides autocomplete suggestions tailored to emphasize stereotypical, counter-stereotypical, or neutral traits in candidate evaluations.
- Novelty:
- Direct manipulation of descriptive language to address gender stereotypes.
- Evaluation of downstream effects on perceptions, hiring decisions, and salary recommendations.
- Examination of potential backlash effects, such as reduced likability of counter-stereotypical candidates.
- Procedure and key techniques:
- Participants (N = 672) evaluated résumés for a male ("John") and a female ("Jennifer") candidate for a financial analyst role.
- AI suggestions for Jennifer varied by condition: stereotypical (warmth-oriented), counter-stereotypical (competence-oriented), or neutral.
- Outcomes measured included written evaluations, trait impressions, hiring decisions, salary offers, and affiliative judgments.
- Language manipulation was achieved using GPT-4-powered autocomplete suggestions, with prompts engineered to align with the experimental conditions.
Results
- Concrete findings:
- Counter-stereotypical suggestions increased Jennifer's perceived competence and eliminated the salary gap between her and John.
- Jennifer was offered similar salaries to John in the counter-stereotypical condition ($64,446 vs. $64,802, p =.278), compared to significant gaps in control and stereotypical conditions.
- Counter-stereotypical suggestions reduced Jennifer's perceived warmth and likability as a colleague.
- Advantage over baselines:
- Counter-stereotypical suggestions reduced salary disparities, a result not observed in the control or stereotypical conditions.
- Participants used more competence-related language when evaluating Jennifer in the counter-stereotypical condition.
- Experiments / evaluation:
- Participants wrote evaluations with AI assistance and completed trait ratings, hiring decisions, and salary recommendations.
- Gender bias in language was quantified using dictionary-based analysis, cosine similarity metrics, and GPT-4 evaluations.
- No significant changes in hiring decisions were observed, though Jennifer was chosen slightly more often in the counter-stereotypical condition (44.6% vs. 41.5% in control, p >.05).
- Limitations and future work:
- Hiring decisions remained resistant to change, potentially due to role congruity and backlash effects.
- Study focused on a single job type and binary gender; future work should explore broader contexts, longer-term effects, and intersectional identities.
- Ethical considerations include transparency, user agency, and potential cultural homogenization.
Summary
This study demonstrates that AI writing assistants can reduce gender bias in hiring evaluations by subtly altering descriptive language. Counter-stereotypical suggestions increased perceptions of competence and eliminated salary disparities for female candidates but reduced their perceived warmth and likability. While hiring decisions showed limited change, the results highlight the potential of language-based interventions to address stereotypes in professional contexts. Future research should explore broader applications, ethical considerations, and long-term impacts of such tools.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
Can AI Be a Moral Victim? The Role of Moral Patiency and Ownership Perceptions in Ethical Judgments of Using AI-Generated Content
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 71%
"I think this is fair": Uncovering the Complexities of Stakeholder Decision-Making in AI Fairness Assessment
CHI '26· AI Ethics, Fairness & Accountability +2
- 67%
A Canary in the AI Coal Mine: American Jews May Be Disproportionately Harmed by Intellectual Property Dispossession in Large Language Model Training
CHI '24· AI Ethics, Fairness & Accountability +1
- 67%
AI Mismatches: Identifying Potential Algorithmic Harms Before AI Development
CHI '25· AI Ethics, Fairness & Accountability +1
Based on Jaccard similarity of research subtopics & professions (≥60%)