It's Trying Too Hard To Look Real: Deepfake Moderation Mistakes and Identity-Based Bias

Privacy by Design & User ControlPrivacy Perception & Decision-MakingDeepfake & Synthetic Media DetectionPrivacy Policy MakersContent Governance & Platform Compliance TeamsSociologists & Anthropologists

Title of the Paper

It’s Trying Too Hard to Look Real: Deepfake Moderation Mistakes and Identity-Based Bias

Paper Information

  • Field of Study: Deepfake content, moderation errors, and identity-based bias
  • Keywords: Deepfake, content moderation, bias, AI mental models, human judgment

Research Background and Problem

  • Identified Problems or Challenges:
    • Social media platforms need to distinguish real users from fake users generated by deepfake technology, but bias leads to misjudgments (e.g., mistaking real accounts for fake ones), a problem that has not been systematically studied in depth.
    • Erroneous identity moderation can cause unfair harm to certain users, particularly those of different gender or racial identities.
  • Significance:
    • Deepfakes are causing increasing economic and emotional harm to users, as well as eroding overall user trust and value on platforms, making this a growing societal issue.
  • Research Motivation and Related Work:
    • Advances in deepfake technology have made it easier to generate realistic fake social media content.
    • Existing research has found significant bias in automated content classification algorithms across different identity groups, but similar analyses of human moderators’ biases are lacking.

Solution

  • Methods or Solutions:
    • Experimental design to analyze human content moderators’ misjudgment of accounts categorized by race and gender.
    • Investigate the impact of account identity information, whether moderators share the same identity as the account, and the specific content displayed by the account (images, names, or text) on misjudgment rates.
  • Innovations:
    • The first systematic analysis of identity bias in manual content moderation errors, exploring the mental models behind human moderation decisions.
    • Provides design recommendations and future directions to mitigate bias.
  • Implementation Steps:
    1. Collect real LinkedIn profiles, totaling 160 accounts.
    2. Recruit 695 participants as moderators to assess the authenticity of the accounts.
    3. Conduct experiments by controlling the conditions of account content display (name and photo only, text only, or all information).
    4. Use quantitative and qualitative analysis to understand the psychological decision-making processes of moderators.

Research Findings

  • Specific Findings:
    • Conclusions on Misjudgment: Moderators’ misjudgments vary significantly based on the racial or gender identity of the accounts, with these variations being more pronounced when less account information is provided.
    • Mental Models: Moderators rely on the following mental models when judging account authenticity:
      1. Worldview or social stereotypes (e.g., occupational expectations).
      2. Speculations about AI functionality (e.g., detecting flawed textures or overly perfect features).
      3. Understanding of attacker strategies (e.g., high-status or ambiguous accounts).
      4. Unarticulated intuition.
  • Advantages:
    • Clearly identifies specific manifestations of identity bias and reveals root causes through quantitative and qualitative analysis.
    • Offers strategies to reduce content moderation bias, including interface design, moderation team composition, and training for safety.
  • Experimental or Evaluation Results:
    • When displaying photos and names, accounts of Black individuals (both male and female) had lower misjudgment rates compared to White accounts. However, when only text was displayed, misjudgment rates were not significantly different across identities.
    • Moderators were less likely to misjudge accounts when they shared the same identity (in-group).
  • Limitations and Future Directions:
    • Limitations include: the study only covered specific racial and gender identities, and did not account for broader identity differences.
    • Future directions include expanding research to cover more identities and professions, developing anti-bias training, and creating more intelligent algorithmic tools.

Conclusion

Through innovative experiments and analytical methods, this study thoroughly explores the issue of identity bias in deepfake content moderation and provides targeted recommendations. It offers valuable insights for platform design, moderation team composition, and understanding user mental models, while urging attention to implicit bias in AI systems and their broader societal impacts.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147959/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3641999
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Privacy by Design & User Control, Privacy Perception & Decision-Making, Deepfake & Synthetic Media Detection
work
Professions
Privacy Policy Makers, Content Governance & Platform Compliance Teams, Sociologists & Anthropologists
article
Content Status
Full text indexed
hub
Related Papers
6 related papers