How Generative AI Empowers Attackers and Defenders Across the Trust & Safety Landscape

Generative AI (Text, Image, Music, Video)AI Ethics, Fairness & AccountabilityPrivacy by Design & User ControlDeepfake & Synthetic Media DetectionOnline Harassment & Counter-ToolsCybersecurity EngineersPrivacy Policy MakersContent Governance & Platform Compliance TeamsHCI Researchers

Paper Title

How Generative AI Empowers Attackers and Defenders Across the Trust & Safety Landscape

Publication Info

  • Topic area: The dual impact of Generative AI on Trust & Safety, focusing on its use by attackers and defenders.
  • Keywords: Generative AI, Trust & Safety, content moderation, cybercrime, misinformation, scams, violent extremism, child safety, hate speech, counternarratives.

Background and Problem

  • Problem / challenge: Generative AI (GenAI) is transforming the Trust & Safety landscape by enabling attackers to scale and enhance harmful activities while defenders struggle to adapt existing systems to counter these threats.
  • Significance: The rapid evolution of GenAI poses significant risks to online safety, including the proliferation of harmful content, scams, and misinformation, while also offering potential tools for defense and mitigation.
  • Motivation and related work: Prior research has focused on technical safeguards for AI misuse but has not adequately addressed the evolving sociotechnical dynamics between attackers and defenders. This paper builds on existing work by engaging domain experts to explore GenAI's specific impacts across five key Trust & Safety domains.

Solution

  • Proposed approach: A qualitative study involving participatory workshops with 43 Trust & Safety experts across five domains: child safety, election integrity, hate and harassment, scams, and violent extremism.
  • Novelty:
    1. A domain-specific analysis of how GenAI empowers both attackers and defenders.
    2. Identification of structural differences and practical implications for Trust & Safety operations.
    3. A strategic framework for leveraging GenAI in defense while addressing its risks.
    4. Insights into the co-evolution of attacker and defender capabilities in the GenAI era.
  • Procedure and key techniques:
    • Conducted six-hour participatory workshops in Asia, Europe, and North America.
    • Participants included experts from civil society, industry, and academia.
    • Activities included memorable experience cards, change cards, and futuring stories to explore GenAI's impact.
    • Data was analyzed through reflexive thematic analysis to identify cross-domain and domain-specific themes.

Results

  • Concrete findings:
    • GenAI enables attackers to scale harmful content, lower barriers to entry, and create sophisticated, personalized attacks.
    • Defenders can leverage GenAI for content moderation, investigations, counternarratives, and user support.
    • GenAI's dual-use nature creates an "arms race" between attackers and defenders.
  • Advantage over baselines:
    • GenAI offers defenders new tools to counter longstanding and emerging threats, such as automating content moderation and improving investigation workflows.
    • However, attackers currently exploit GenAI more effectively, outpacing existing defenses.
  • Experiments / evaluation:
    • Workshops with 43 experts across five domains provided qualitative insights into GenAI's impact.
    • Participants highlighted domain-specific challenges and opportunities, such as the need for structural changes in child safety and nuanced policy for election integrity.
  • Limitations and future work:
    • Findings are contextually bounded to the specific domains and participants studied.
    • Rapidly evolving GenAI technologies may shift expert perspectives over time.
    • Future work should explore cross-sector collaboration and develop standardized benchmarks for GenAI applications in Trust & Safety.

Summary

This paper explores how Generative AI is reshaping the Trust & Safety landscape by empowering both attackers and defenders. Through a qualitative study with 43 experts across five domains, the authors identify how GenAI increases the scale and sophistication of attacks while offering new defensive capabilities. Key contributions include insights into domain-specific challenges, the co-evolution of attacker and defender strategies, and the need for structural changes in Trust & Safety operations. The findings underscore the urgency of leveraging GenAI responsibly to create safer online environments while addressing its inherent risks.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222347/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791363
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), AI Ethics, Fairness & Accountability, Privacy by Design & User Control, Deepfake & Synthetic Media Detection
work
Professions
Cybersecurity Engineers, Privacy Policy Makers, Content Governance & Platform Compliance Teams, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers