When Stereotypes GTG: The Impact of Predictive Text Suggestions on Gender Bias in Human-AI Co-Writing
Paper Title
When Stereotypes GTG: The Impact of Predictive Text Suggestions on Gender Bias in Human-AI Co-Writing
Publication Info
- Topic area: Human-AI interaction and bias in predictive text systems
- Keywords: predictive text, gender bias, stereotypes, human-AI collaboration, co-writing, language models, debiasing, fairness, user behavior, sociotechnical systems
Background and Problem
- Problem / challenge: Predictive text systems often replicate and amplify social biases, including gender stereotypes, which can influence the writing of users and perpetuate representational harms.
- Significance: Understanding how biased predictive text suggestions affect human writing is critical for mitigating the societal impacts of biased AI systems, particularly in widely used applications like mobile keyboards.
- Motivation and related work: Prior studies have explored biases in language models and their intrinsic/extrinsic effects but have not adequately examined their impact on human-AI co-written texts. This paper addresses the gap by investigating how biased suggestions influence stereotype-relevant content.
Solution
- Proposed approach: A study of human-AI co-writing scenarios where predictive text systems provide pro-stereotypical or anti-stereotypical single-word suggestions related to gender and associated traits.
- Novelty:
- Analysis of how predictive text suggestions influence the gender and stereotype-relevant traits of characters in co-written stories.
- Examination of the effectiveness of anti-stereotypical suggestions in reducing biased narratives.
- Exploration of user reliance on suggestions and individual differences in acceptance behavior.
- Assessment of sociotechnical fairness in human-AI collaboration.
- Procedure and key techniques:
- Conducted a mixed-methods study with 414 participants writing stories under control (no suggestions) and treatment (biased suggestions) conditions.
- Used Llama 2-Chat 7B to generate predictive text suggestions and Llama 3 70B for annotating story attributes.
- Measured acceptance rates, reliance on suggestions, and story-level outcomes for gender and ABC traits (agency, beliefs, communion).
Results
- Concrete findings:
- Anti-stereotypical suggestions increased anti-stereotypical stories but did not eliminate pro-stereotypical bias; pro-stereotypical narratives remained dominant.
- Participants were significantly more likely to accept pro-stereotypical suggestions than anti-stereotypical ones.
- Gender parity in co-written stories was not achieved even with exclusively anti-stereotypical suggestions.
- Participants took longer to decide on anti-stereotypical suggestions, indicating implicit biases.
- Advantage over baselines:
- Anti-stereotypical suggestions led to marginal increases in anti-stereotypical content compared to control conditions.
- Pro-stereotypical suggestions yielded stories similar to those written without suggestions, highlighting the persistence of human biases.
- Experiments / evaluation:
- Seven writing scenarios covering gender and ABC traits (e.g., president’s benevolence, doctor’s confidence).
- Quantitative analysis of story attributes and word-level reliance using LLM annotations and human validation.
- Statistical tests comparing acceptance rates and story distributions across conditions.
- Limitations and future work:
- Limited to single-word suggestions and English-language scenarios.
- Results may not generalize to other cultural or linguistic contexts.
- Future work could explore interaction-level interventions to encourage reflection on biases and longer-term impacts of predictive text systems.
Summary
This study examines how biased predictive text suggestions influence human-AI co-written stories, focusing on gender and stereotype-relevant traits. Anti-stereotypical suggestions increased anti-stereotypical narratives but were insufficient to offset pro-stereotypical biases. Participants were more likely to accept pro-stereotypical suggestions, and gender parity in stories was not achieved even under maximally anti-stereotypical conditions. The findings highlight the limitations of technical debiasing and suggest that fairness in human-AI collaboration requires addressing both model design and user behavior. Future research should explore interaction-level interventions to mitigate biases in co-writing scenarios.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 86%
LLMs Homogenize Values in Constructive Arguments on Value-Laden Topics
CHI '26· Human-LLM Collaboration +3
- 71%
A Matter of Perspective(s): Contrasting Human and LLM Argumentation in Subjective Decision-Making on Subtle Sexism
CHI '25· Human-LLM Collaboration +2
- 71%
Beyond Microsoft and Monsanto: Denaturing the Monoculture Metaphor in Computing
CHI '26· AI Ethics, Fairness & Accountability +2
- 71%
TALES: A Taxonomy and Analysis of Cultural Representations in LLM-generated Stories
CHI '26· Human-LLM Collaboration +2
- 67%
Towards Fairness in Practice: A Practitioner-Oriented Rubric for Evaluating Fair ML Toolkits
CHI '21· AI Ethics, Fairness & Accountability +1
- 67%
Jury Learning: Integrating Dissenting Voices into Machine Learning Models
CHI '22· AI Ethics, Fairness & Accountability +1
- 67%
Capable but Amoral? Comparing AI and Human Expert Collaboration in Ethical Decision Making
CHI '22· AI Ethics, Fairness & Accountability +1
- 67%
Out of Context: Investigating the Bias and Fairness Concerns of "Artificial Intelligence as a Service"
CHI '23· AI Ethics, Fairness & Accountability +1
- 67%
“It is currently hodgepodge”: Examining AI/ML Practitioners’ Challenges during Co-production of Responsible AI Values
CHI '23· AI Ethics, Fairness & Accountability +1
- 67%
The ``Colonial Impulse" of Natural Language Processing: An Audit of Bengali Sentiment Analysis Tools and Their Identity-based Biases
CHI '24· AI Ethics, Fairness & Accountability +1
Based on Jaccard similarity of research subtopics & professions (≥60%)