How Generative AI Empowers Attackers and Defenders Across the Trust & Safety Landscape
Authors
Paper Title
How Generative AI Empowers Attackers and Defenders Across the Trust & Safety Landscape
Publication Info
- Topic area: The dual impact of Generative AI on Trust & Safety, focusing on its use by attackers and defenders.
- Keywords: Generative AI, Trust & Safety, content moderation, cybercrime, misinformation, scams, violent extremism, child safety, hate speech, counternarratives.
Background and Problem
- Problem / challenge: Generative AI (GenAI) is transforming the Trust & Safety landscape by enabling attackers to scale and enhance harmful activities while defenders struggle to adapt existing systems to counter these threats.
- Significance: The rapid evolution of GenAI poses significant risks to online safety, including the proliferation of harmful content, scams, and misinformation, while also offering potential tools for defense and mitigation.
- Motivation and related work: Prior research has focused on technical safeguards for AI misuse but has not adequately addressed the evolving sociotechnical dynamics between attackers and defenders. This paper builds on existing work by engaging domain experts to explore GenAI's specific impacts across five key Trust & Safety domains.
Solution
- Proposed approach: A qualitative study involving participatory workshops with 43 Trust & Safety experts across five domains: child safety, election integrity, hate and harassment, scams, and violent extremism.
- Novelty:
- A domain-specific analysis of how GenAI empowers both attackers and defenders.
- Identification of structural differences and practical implications for Trust & Safety operations.
- A strategic framework for leveraging GenAI in defense while addressing its risks.
- Insights into the co-evolution of attacker and defender capabilities in the GenAI era.
- Procedure and key techniques:
- Conducted six-hour participatory workshops in Asia, Europe, and North America.
- Participants included experts from civil society, industry, and academia.
- Activities included memorable experience cards, change cards, and futuring stories to explore GenAI's impact.
- Data was analyzed through reflexive thematic analysis to identify cross-domain and domain-specific themes.
Results
- Concrete findings:
- GenAI enables attackers to scale harmful content, lower barriers to entry, and create sophisticated, personalized attacks.
- Defenders can leverage GenAI for content moderation, investigations, counternarratives, and user support.
- GenAI's dual-use nature creates an "arms race" between attackers and defenders.
- Advantage over baselines:
- GenAI offers defenders new tools to counter longstanding and emerging threats, such as automating content moderation and improving investigation workflows.
- However, attackers currently exploit GenAI more effectively, outpacing existing defenses.
- Experiments / evaluation:
- Workshops with 43 experts across five domains provided qualitative insights into GenAI's impact.
- Participants highlighted domain-specific challenges and opportunities, such as the need for structural changes in child safety and nuanced policy for election integrity.
- Limitations and future work:
- Findings are contextually bounded to the specific domains and participants studied.
- Rapidly evolving GenAI technologies may shift expert perspectives over time.
- Future work should explore cross-sector collaboration and develop standardized benchmarks for GenAI applications in Trust & Safety.
Summary
This paper explores how Generative AI is reshaping the Trust & Safety landscape by empowering both attackers and defenders. Through a qualitative study with 43 experts across five domains, the authors identify how GenAI increases the scale and sophistication of attacks while offering new defensive capabilities. Key contributions include insights into domain-specific challenges, the co-evolution of attacker and defender strategies, and the need for structural changes in Trust & Safety operations. The findings underscore the urgency of leveraging GenAI responsibly to create safer online environments while addressing its inherent risks.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)