Gold Standard or Gold-Plated? Human Practices of Triple Verification in CSAM Takedown
Honorable MentionAuthors
Paper Title
Gold Standard or Gold-Plated? Human Practices of Triple Verification in CSAM Takedown
Publication Info
- Topic area: Verification practices for CSAM classification and takedown.
- Keywords: CSAM, triple verification, inter-rater reliability, content moderation, hash databases, blind vs. non-blind review, psychological burden, AI in moderation, series recognition, Digital Services Act.
Background and Problem
- Problem / challenge: Current verification procedures for CSAM classification, particularly triple verification, are poorly understood. There is limited research on how these procedures are implemented, their impact on accuracy and efficiency, and the challenges faced by experts.
- Significance: Accurate CSAM classification is critical for issuing Notice-and-Takedown (NTD) requests, maintaining hash databases, and ensuring legal compliance. Misclassification can lead to severe consequences, including continued victimization or wrongful inclusion in databases.
- Motivation and related work: Previous studies have focused on CSAM detection tools and user perceptions of automated scanning but have not explored verification procedures or their operationalization. Forensic research has examined inter-rater reliability but lacks insights into real-world workflows, professional reviewers, and varying conditions like blind vs. non-blind voting.
Solution
- Proposed approach: A mixed-methods study combining expert interviews, an inter-rater reliability experiment, and a focus group to investigate verification practices, expert perceptions, and the impact of voting conditions on CSAM classification.
- Novelty:
- Mapping diverse verification procedures across organizations.
- Empirically assessing inter-rater reliability under blind and non-blind conditions.
- Identifying key challenges in expert disagreements, including series recognition and age estimation.
- Proposing adaptive verification practices to balance accuracy, efficiency, and reviewer well-being.
- Procedure and key techniques:
- Conducted 14 semi-structured interviews with experts from seven organizations to understand workflows and perceptions.
- Designed an inter-rater reliability experiment with Dutch National Police experts, analyzing 2,031 images/videos under blind and non-blind conditions.
- Held a focus group to explore reasons behind disagreements and contextualize experimental findings.
Results
- Concrete findings:
- Non-blind conditions increased agreement (Cohen’s Kappa: 0.893 vs. 0.670 in blind conditions).
- Videos were classified more reliably than images (Kappa: 0.812 for videos vs. 0.601 for images in blind conditions).
- Triple verification revealed disagreements in 8% of cases compared to double verification.
- Advantage over baselines:
- Triple verification surfaced residual uncertainties and improved accuracy in ambiguous cases.
- Non-blind workflows enhanced convergence but introduced potential bias.
- Experiments / evaluation:
- Dataset: 2,031 CSAM-related items reviewed by Dutch National Police experts.
- Metrics: Cohen’s Kappa for inter-rater reliability, raw agreement percentages.
- Conditions: Blind vs. non-blind voting, impact of voting order, file type differences.
- Limitations and future work:
- Limited diversity in interview participants (e.g., no response from NCMEC).
- Findings are specific to Dutch legal and operational contexts.
- Future work should explore AI’s role in reducing workload and handling synthetic CSAM.
Summary
This study investigates the practices and challenges of CSAM verification, focusing on triple verification as a safeguard against misclassification. Through interviews, experiments, and focus groups, the authors reveal that verification practices vary widely, with non-blind workflows increasing agreement but raising concerns about bias. Key challenges include series recognition, age estimation, and the psychological toll on reviewers. The findings highlight the trade-offs between accuracy, efficiency, and reviewer well-being, suggesting adaptive verification practices and cautious integration of AI to support human oversight. These insights contribute to improving CSAM governance and ensuring ethical, scalable verification processes.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 60%
Upstanding by Design: Bystander Intervention in Cyberbullying
CHI '18· Online Harassment & Counter-Tools +1
- 60%
Managing Deviant Behavior in Online Communities III
CHI '18· Online Harassment & Counter-Tools +1
- 60%
Characterizing Twitter Users Who Engage in Adversarial Interactions against Political Candidates
CHI '20· Online Harassment & Counter-Tools +1
- 60%
In Suspense About Suspensions? The Relative Effectiveness of Suspension Durations on a Popular Social Platform
CHI '25· Online Harassment & Counter-Tools +1
Based on Jaccard similarity of research subtopics & professions (≥60%)