Identifying, Explaining, and Correcting Ableist Language with AI

Human-LLM CollaborationAI Ethics, Fairness & AccountabilityCognitive Impairment & Neurodiversity (Autism, ADHD, Dyslexia)Inclusive DesignPsychiatrists & PsychotherapistsAssistive Technology SpecialistsHCI Researchers

Paper Title

Identifying, Explaining, and Correcting Ableist Language with AI

Publication Info

  • Topic area: AI-assisted detection and correction of ableist language in text.
  • Keywords: Ableism, AI annotations, inclusive language, disability justice, GPT-4o, bias detection, narrative analysis, human-AI comparison, cultural competence, writing tools.

Background and Problem

  • Problem / challenge: Subtle ableist language remains pervasive and difficult to recognize, often overlooked by general hate-speech tools and communicators. Current frameworks lack adaptability to nuanced contexts and rely heavily on human expertise, which is not scalable.
  • Significance: Addressing ableist language is crucial for fostering inclusive communication and reducing harm to disabled communities in everyday writing, journalism, and professional contexts.
  • Motivation and related work: Prior research highlights AI’s tendency to reproduce ableist norms and the challenges of detecting implicit harms. While automated tools have shown promise in addressing other biases (e.g., racism, sexism), their application to ableism remains underexplored. This paper builds on existing studies to evaluate AI’s role in identifying, explaining, and correcting ableist language, comparing its performance to human annotations.

Solution

  • Proposed approach: A two-part study evaluating GPT-4o’s ability to annotate ableist language in fictional narratives, comparing AI-generated annotations with human crowdsourced annotations from the disability community.
  • Novelty:
    1. Creation of a first-of-its-kind dataset of nuanced ableism annotations rooted in lived experience.
    2. Empirical comparison of AI- and human-generated annotations, highlighting tradeoffs between consistency, clarity, and cultural depth.
    3. Design guidelines for AI systems and writing tools addressing culturally sensitive language.
  • Procedure and key techniques:
    • Study 1: Generate fictional stories containing subtle ableism using GPT-4o, validated by pilot participants from disability groups. Collect human annotations identifying, explaining, and correcting ableist language at sentence and passage levels.
    • Study 2: Compare AI-generated annotations (produced using chained prompting) with human baseline annotations synthesized from Study 1. Survey 108 participants from the disability community to evaluate annotation quality, agreeableness, and preferences.

Results

  • Concrete findings:
    • Sentence-level agreement with AI and human annotations averaged 72.3%.
    • Participants preferred AI annotations overall (43.9%) over human annotations (23.4%), citing clarity, consistency, and accessible formatting.
    • Human annotations were valued for cultural grounding and advocacy-oriented framing but criticized for dense phrasing and inconsistent logic between explanations and corrections.
  • Advantage over baselines:
    • AI annotations were praised for producing concise, neutral, and actionable feedback, while human annotations excelled in emotional resonance and cultural critique.
  • Experiments / evaluation:
    • Study 1 involved 110 participants annotating AI-generated fictional stories across seven disability categories. Study 2 surveyed 108 participants to compare AI and human annotations using Likert scales and qualitative feedback.
    • Metrics included agreement rates, preference rankings, and thematic coding of participant feedback.
  • Limitations and future work:
    • Findings may not generalize to nonfiction or professional writing contexts.
    • Participant pool was limited to college-educated individuals with DEI training, not fully representative of the broader disability community.
    • Future work should expand datasets, test long-term educational impacts, and explore applications across other forms of bias.

Summary

This paper investigates the role of AI in identifying, explaining, and correcting ableist language, comparing GPT-4o annotations with human annotations from the disability community. Results show comparable accuracy but distinct strengths: AI annotations excel in clarity and consistency, while human annotations provide cultural depth and advocacy-oriented framing. Participants preferred AI annotations overall, highlighting their potential as scalable tools for bias education. Contributions include a nuanced ableism dataset, empirical insights into human-AI annotation tradeoffs, and design guidelines for inclusive writing tools. Future work aims to expand datasets, test broader applications, and refine AI systems to balance clarity with cultural sensitivity.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222740/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790533
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Human-LLM Collaboration, AI Ethics, Fairness & Accountability, Cognitive Impairment & Neurodiversity (Autism, ADHD, Dyslexia), Inclusive Design
work
Professions
Psychiatrists & Psychotherapists, Assistive Technology Specialists, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers