When Stereotypes GTG: The Impact of Predictive Text Suggestions on Gender Bias in Human-AI Co-Writing

Human-LLM CollaborationAI Ethics, Fairness & AccountabilityAlgorithmic Fairness & BiasAI/ML Researchers & EngineersHCI ResearchersSociologists & Anthropologists

Paper Title

When Stereotypes GTG: The Impact of Predictive Text Suggestions on Gender Bias in Human-AI Co-Writing

Publication Info

  • Topic area: Human-AI interaction and bias in predictive text systems
  • Keywords: predictive text, gender bias, stereotypes, human-AI collaboration, co-writing, language models, debiasing, fairness, user behavior, sociotechnical systems

Background and Problem

  • Problem / challenge: Predictive text systems often replicate and amplify social biases, including gender stereotypes, which can influence the writing of users and perpetuate representational harms.
  • Significance: Understanding how biased predictive text suggestions affect human writing is critical for mitigating the societal impacts of biased AI systems, particularly in widely used applications like mobile keyboards.
  • Motivation and related work: Prior studies have explored biases in language models and their intrinsic/extrinsic effects but have not adequately examined their impact on human-AI co-written texts. This paper addresses the gap by investigating how biased suggestions influence stereotype-relevant content.

Solution

  • Proposed approach: A study of human-AI co-writing scenarios where predictive text systems provide pro-stereotypical or anti-stereotypical single-word suggestions related to gender and associated traits.
  • Novelty:
    1. Analysis of how predictive text suggestions influence the gender and stereotype-relevant traits of characters in co-written stories.
    2. Examination of the effectiveness of anti-stereotypical suggestions in reducing biased narratives.
    3. Exploration of user reliance on suggestions and individual differences in acceptance behavior.
    4. Assessment of sociotechnical fairness in human-AI collaboration.
  • Procedure and key techniques:
    • Conducted a mixed-methods study with 414 participants writing stories under control (no suggestions) and treatment (biased suggestions) conditions.
    • Used Llama 2-Chat 7B to generate predictive text suggestions and Llama 3 70B for annotating story attributes.
    • Measured acceptance rates, reliance on suggestions, and story-level outcomes for gender and ABC traits (agency, beliefs, communion).

Results

  • Concrete findings:
    • Anti-stereotypical suggestions increased anti-stereotypical stories but did not eliminate pro-stereotypical bias; pro-stereotypical narratives remained dominant.
    • Participants were significantly more likely to accept pro-stereotypical suggestions than anti-stereotypical ones.
    • Gender parity in co-written stories was not achieved even with exclusively anti-stereotypical suggestions.
    • Participants took longer to decide on anti-stereotypical suggestions, indicating implicit biases.
  • Advantage over baselines:
    • Anti-stereotypical suggestions led to marginal increases in anti-stereotypical content compared to control conditions.
    • Pro-stereotypical suggestions yielded stories similar to those written without suggestions, highlighting the persistence of human biases.
  • Experiments / evaluation:
    • Seven writing scenarios covering gender and ABC traits (e.g., president’s benevolence, doctor’s confidence).
    • Quantitative analysis of story attributes and word-level reliance using LLM annotations and human validation.
    • Statistical tests comparing acceptance rates and story distributions across conditions.
  • Limitations and future work:
    • Limited to single-word suggestions and English-language scenarios.
    • Results may not generalize to other cultural or linguistic contexts.
    • Future work could explore interaction-level interventions to encourage reflection on biases and longer-term impacts of predictive text systems.

Summary

This study examines how biased predictive text suggestions influence human-AI co-written stories, focusing on gender and stereotype-relevant traits. Anti-stereotypical suggestions increased anti-stereotypical narratives but were insufficient to offset pro-stereotypical biases. Participants were more likely to accept pro-stereotypical suggestions, and gender parity in stories was not achieved even under maximally anti-stereotypical conditions. The findings highlight the limitations of technical debiasing and suggest that fairness in human-AI collaboration requires addressing both model design and user behavior. Future research should explore interaction-level interventions to mitigate biases in co-writing scenarios.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222161/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790733
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
2 authors
sell
Subtopics
Human-LLM Collaboration, AI Ethics, Fairness & Accountability, Algorithmic Fairness & Bias
work
Professions
AI/ML Researchers & Engineers, HCI Researchers, Sociologists & Anthropologists
article
Content Status
Full text indexed
hub
Related Papers
10 related papers