Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations
Authors
AI Ethics, Fairness & AccountabilityAlgorithmic Transparency & AuditabilityDark Patterns RecognitionContent Moderation & Platform GovernancePrivacy Policy MakersContent Governance & Platform Compliance TeamsHCI Researchers
Research Background and Issues
- Identified Problems or Challenges: Commercial content moderation APIs exhibit two significant failures when addressing online hate speech: over-moderation (mislabeling legitimate content as hate speech) and under-moderation (failing to detect actual hate speech). These issues are particularly pronounced for content using identity labels (e.g., "Black" or "gay"). There are performance disparities in moderating content targeting specific groups and language variants, disproportionately affecting minority groups.
- Significance: False labeling of hate speech suppresses legitimate expression, undermining the diversity of public discourse, while undetected hate speech exacerbates toxicity on social media and harms targeted groups. These issues threaten equitable communication environments for minority groups such as LGBTQIA+, Jewish, and Muslim communities, impacting user trust.
- Related Work: Numerous studies have explored issues in open-source NLP models, including their vilification of certain target groups and failure to detect implicit hate speech. However, systematic evaluations of commercial APIs remain scarce, as the "black-box" nature of these APIs limits transparency in related research.
Solution
- Proposed Method or Solution: This paper introduces a novel framework for auditing commercial content moderation APIs, focusing on five commonly used services (e.g., OpenAI, Google, Amazon). It evaluates their performance by analyzing over five million queries and four benchmark datasets.
- Innovations:
- Provides the first reproducible framework for auditing "black-box" NLP models.
- Develops three experiments (performance comparison, counterfactual fairness analysis, and SHAP explainability analysis) to deeply investigate functional errors and bias sources in each API.
- Goes beyond simple error rate statistics by incorporating corpus analysis to explain the models' linguistic preferences (e.g., over-reliance on specific group identity terms).
- Implementation Steps:
- Utilize five commercial content moderation APIs to process over 5 million queries and collect data.
- Employ four benchmark datasets (e.g., ToxiGen and HateXplain) to evaluate hate speech categories and performance across target groups.
- Design three experiments: overall and target group performance analysis, perturbation sensitivity analysis to detect group biases, and SHAP-based analysis to interpret model preferences for specific vocabulary.
Research Findings
- Specific Findings: All APIs were found to predict hate speech based on identity labels (e.g., "Black" or "gay"), leading to:
- Under-moderation of content targeting LGBTQIA+ groups and implicit hate speech.
- Over-moderation of reclaimed slurs (e.g., the Black community's reclamation of "n*gger") and counter-speech.
- Strengths: OpenAI and Amazon demonstrated slightly better performance compared to other services. These services performed relatively well on explicit hate speech but still struggled with detecting implicit hate speech.
- Experimental or Evaluation Results: Bulk analysis revealed that Google Natural Language API performed the worst for minority groups, frequently deleting legitimate content (with a false positive rate as high as 99%). SHAP explanations showed that many models were overly sensitive to specific vocabulary associated with minority groups, leading to misclassification.
- Limitations and Future Directions:
- Limitations: The black-box nature of APIs prevents access to model weights or training data; high query costs limit the scale of audits; existing benchmark dataset labels may not be entirely accurate.
- Future Directions: Future research should expand to non-English content moderation, explore performance on other content types (e.g., images, audio), and develop more efficient auditing methods to reduce costs. Additionally, clearer requirements for multilingual support and nuanced understanding within APIs should be established.
Conclusion
This study, through a systematic framework and experiments, reveals the performance deficiencies of commercial content moderation APIs and their negative impact on specific groups, highlighting the importance of transparency and the necessity of auditing. The paper calls on content moderation service providers to optimize algorithm design and provide greater transparency, while recommending policymakers to enhance legal support and accountability for third-party audits. This work provides theoretical foundations and practical tools for the future development of content moderation systems.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- What functional deficiencies do commercial content moderation APIs have in handling hate speech?Category: Fairness, Bias, and Cultural Adaptation in Online Content ModerationSimilar questionsarrow_forward
- Do these APIs show bias when moderating content across target groups and language varieties?Category: Fairness, Bias, and Cultural Adaptation in Online Content ModerationSimilar questionsarrow_forward
- How can a reproducible framework be designed to audit black-box NLP models for performance and fairness?Category: Fairness, Bias, and Cultural Adaptation in Online Content ModerationSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Legitimate expression by minority groups is often wrongly removed while hate speech is missed by content moderation.Category: Fairness, Bias, and Cultural Adaptation in Online Content ModerationSimilar questionsarrow_forward
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713998
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
AI Ethics, Fairness & Accountability, Algorithmic Transparency & Auditability, Dark Patterns Recognition, Content Moderation & Platform Governance
work
Professions
Privacy Policy Makers, Content Governance & Platform Compliance Teams, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers