Gold Standard or Gold-Plated? Human Practices of Triple Verification in CSAM Takedown

Honorable Mention
Online Harassment & Counter-ToolsContent Moderation & Platform GovernancePolice & Emergency Service PersonnelContent Governance & Platform Compliance Teams

Paper Title

Gold Standard or Gold-Plated? Human Practices of Triple Verification in CSAM Takedown

Publication Info

  • Topic area: Verification practices for CSAM classification and takedown.
  • Keywords: CSAM, triple verification, inter-rater reliability, content moderation, hash databases, blind vs. non-blind review, psychological burden, AI in moderation, series recognition, Digital Services Act.

Background and Problem

  • Problem / challenge: Current verification procedures for CSAM classification, particularly triple verification, are poorly understood. There is limited research on how these procedures are implemented, their impact on accuracy and efficiency, and the challenges faced by experts.
  • Significance: Accurate CSAM classification is critical for issuing Notice-and-Takedown (NTD) requests, maintaining hash databases, and ensuring legal compliance. Misclassification can lead to severe consequences, including continued victimization or wrongful inclusion in databases.
  • Motivation and related work: Previous studies have focused on CSAM detection tools and user perceptions of automated scanning but have not explored verification procedures or their operationalization. Forensic research has examined inter-rater reliability but lacks insights into real-world workflows, professional reviewers, and varying conditions like blind vs. non-blind voting.

Solution

  • Proposed approach: A mixed-methods study combining expert interviews, an inter-rater reliability experiment, and a focus group to investigate verification practices, expert perceptions, and the impact of voting conditions on CSAM classification.
  • Novelty:
    1. Mapping diverse verification procedures across organizations.
    2. Empirically assessing inter-rater reliability under blind and non-blind conditions.
    3. Identifying key challenges in expert disagreements, including series recognition and age estimation.
    4. Proposing adaptive verification practices to balance accuracy, efficiency, and reviewer well-being.
  • Procedure and key techniques:
    1. Conducted 14 semi-structured interviews with experts from seven organizations to understand workflows and perceptions.
    2. Designed an inter-rater reliability experiment with Dutch National Police experts, analyzing 2,031 images/videos under blind and non-blind conditions.
    3. Held a focus group to explore reasons behind disagreements and contextualize experimental findings.

Results

  • Concrete findings:
    • Non-blind conditions increased agreement (Cohen’s Kappa: 0.893 vs. 0.670 in blind conditions).
    • Videos were classified more reliably than images (Kappa: 0.812 for videos vs. 0.601 for images in blind conditions).
    • Triple verification revealed disagreements in 8% of cases compared to double verification.
  • Advantage over baselines:
    • Triple verification surfaced residual uncertainties and improved accuracy in ambiguous cases.
    • Non-blind workflows enhanced convergence but introduced potential bias.
  • Experiments / evaluation:
    • Dataset: 2,031 CSAM-related items reviewed by Dutch National Police experts.
    • Metrics: Cohen’s Kappa for inter-rater reliability, raw agreement percentages.
    • Conditions: Blind vs. non-blind voting, impact of voting order, file type differences.
  • Limitations and future work:
    • Limited diversity in interview participants (e.g., no response from NCMEC).
    • Findings are specific to Dutch legal and operational contexts.
    • Future work should explore AI’s role in reducing workload and handling synthetic CSAM.

Summary

This study investigates the practices and challenges of CSAM verification, focusing on triple verification as a safeguard against misclassification. Through interviews, experiments, and focus groups, the authors reveal that verification practices vary widely, with non-blind workflows increasing agreement but raising concerns about bias. Key challenges include series recognition, age estimation, and the psychological toll on reviewers. The findings highlight the trade-offs between accuracy, efficiency, and reviewer well-being, suggesting adaptive verification practices and cautious integration of AI to support human oversight. These insights contribute to improving CSAM governance and ensuring ethical, scalable verification processes.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222850/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791039
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
Honorable Mention
group
Authors
3 authors
sell
Subtopics
Online Harassment & Counter-Tools, Content Moderation & Platform Governance
work
Professions
Police & Emergency Service Personnel, Content Governance & Platform Compliance Teams
article
Content Status
Full text indexed
hub
Related Papers
4 related papers