Taking the control back – An adventure in developing personalized content moderation
Honorable MentionAuthors
Paper Title
Taking the control back – An adventure in developing personalized content moderation
Publication Info
- Topic area: Personalized content moderation for combating online harassment.
- Keywords: Online harassment, content moderation, personalized tools, social media, Twitter, automation, API limitations, collaborative moderation, platform design, user empowerment.
Background and Problem
- Problem / challenge: Online platforms fail to provide effective tools for addressing severe, targeted harassment. Current moderation tools are limited, reactive, and often expose victims to further harm.
- Significance: Harassment severely impacts user well-being and social interactions, particularly for women and marginalized groups. Effective moderation is critical for creating safe online spaces.
- Motivation and related work: Existing moderation approaches include automated systems, manual reporting, and community-based tools. However, these methods are often insufficient for addressing persistent, targeted harassment campaigns. This paper builds on prior work by exploring a personalized, automated approach to content moderation.
Solution
- Proposed approach: A personalized, automated, and collaborative anti-harassment system designed to filter, block, and report harassers while minimizing the victim’s exposure to harmful content.
- Novelty:
- Development of a personalized anti-harassment system tailored to a specific harassment campaign.
- Integration of automation and collaboration with friends to enhance reporting and blocking.
- Use of reverse-engineered APIs after official APIs became inaccessible.
- Empirical insights from an autoethnographic study of a sustained harassment campaign.
- Procedure and key techniques:
- Automated detection of harassing accounts based on behavioral patterns.
- Blocking and reporting harassers using a combination of user and supporter accounts.
- Logging and visualization of harassment data for analysis and system refinement.
- Adaptation to platform changes, including reverse-engineering Twitter’s web client.
Results
- Concrete findings:
- The system blocked all harassing accounts within minutes of their first interaction.
- 96% of the 267 harassing accounts were suspended or deleted after automated reporting.
- Over 81,000 reports were submitted, saving an estimated 67.54 hours of manual effort.
- Advantage over baselines:
- Higher suspension rate (96%) compared to typical manual reporting outcomes (e.g., 55% in prior studies).
- Significant reduction in the victim’s exposure to harmful content.
- Experiments / evaluation:
- Analysis of harassment patterns, including account creation and posting behaviors.
- System performance evaluated through logs, visualizations, and reporting outcomes.
- Comparison of moderation features across multiple platforms.
- Limitations and future work:
- System effectiveness is limited to cases of severe, targeted harassment by identifiable accounts.
- Current implementation requires technical expertise; non-technical users may face barriers.
- Future work includes developing user-friendly interfaces and extending the system to address coordinated harassment campaigns.
Summary
This paper presents a personalized, automated anti-harassment system developed in response to a sustained harassment campaign on Twitter. The system effectively blocked and reported harassers, achieving a 96% suspension rate for abusive accounts while minimizing the victim’s exposure to harmful content. By integrating automation, collaboration, and reverse-engineering, the system addresses gaps in existing platform tools. The findings highlight the need for better platform-level designs and user empowerment in combating online harassment. This work provides a foundation for future research and tool development to support victims of severe harassment.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
A Systematic Review of User Experiments on the Effects of Dark Patterns
CHI '26· Dark Patterns Recognition +1
- 71%
Mapping Social Media Dependency: Functional and Psychological Platform Reliance as Mechanisms of Digital Vulnerability
CHI '26· Social Platform Design & User Behavior +2
- 67%
"Okay, whatever": An Evaluation of Cookie Consent Interfaces
CHI '22· Privacy Perception & Decision-Making +1
Based on Jaccard similarity of research subtopics & professions (≥60%)