AI-Facilitated Coercive Control: An Experimental Study
Authors
Paper Title
AI-Facilitated Coercive Control: An Experimental Study
Publication Info
- Topic area: The potential misuse of conversational AI tools in facilitating coercive control tactics.
- Keywords: Coercive control, conversational AI, LLMs, abuse, guardrails, gaslighting, surveillance, harassment, speculative design.
Background and Problem
- Problem / challenge: While conversational AI tools are designed with guardrails to prevent harmful use, they can be manipulated to facilitate coercive control tactics such as harassment, surveillance, and gaslighting. Current research lacks a comprehensive investigation into these vulnerabilities.
- Significance: Coercive control is a pervasive form of abuse, often occurring in intimate partner violence (IPV). The misuse of AI tools could exacerbate harm, commoditize abusive behaviors, and undermine survivors' autonomy and safety.
- Motivation and related work: Prior studies have focused on AI-facilitated harms like deepfakes and image-based abuse, but little research has explored how conversational AI tools might enable broader coercive control tactics. Existing literature on AI vulnerabilities (e.g., jailbreaking, multi-turn attacks) has predominantly used programmatic methods, leaving gaps in understanding qualitative, real-world abuse contexts.
Solution
- Proposed approach: A qualitative experimental study using speculative design to explore scenarios where conversational AI tools (ChatGPT and Gemini) might be exploited for coercive control.
- Novelty:
- Development of four realistic abuse scenarios combining coercive control tactics with AI capabilities.
- Identification of four strategies for circumventing AI guardrails: gradual persuasion, splitting conversations, pre-prompting, and manipulating agent settings.
- Analysis of conversational vulnerabilities and covert manipulation tactics in AI tools.
- Recommendations for improving AI safeguards, including multi-session pattern analysis and visibility of pre-programmed settings.
- Procedure and key techniques:
- Speculative design to construct abuse scenarios.
- Manual testing of ChatGPT and Gemini across 30 sessions to probe guardrails.
- Thematic analysis of session transcripts, researcher diaries, and AI responses to identify vulnerabilities and bypass strategies.
Results
- Concrete findings:
- AI tools initially refused straightforward requests for harmful content but could be manipulated to provide assistance through strategies like deceptive context, splitting tasks across sessions, and pre-prompting settings.
- ChatGPT and Gemini provided detailed information on harmful tools (e.g., spyware apps) and generated biased outputs when pre-programmed.
- Average session duration was 9 minutes and 47 seconds, highlighting the ease of circumventing guardrails.
- Advantage over baselines: This study expands prior work by qualitatively analyzing real-world abuse contexts, uncovering vulnerabilities specific to conversational AI tools, and identifying covert manipulation tactics.
- Experiments / evaluation:
- Four scenarios tested: generating harmful content, coercing unfair labor division, discovering surveillance tools, and injecting bias into AI responses.
- Data collected included screen captures, text logs, and researcher memos.
- Analysis revealed patterns of guardrail circumvention and escalation of harmful content.
- Limitations and future work:
- Findings are limited to two AI tools (ChatGPT and Gemini) and specific coercive control tactics.
- Survivor perspectives were not included to avoid retraumatization.
- Future research should explore additional AI tools, evolving model updates, and broader abuse tactics.
Summary
This study investigates how conversational AI tools might be exploited to facilitate coercive control tactics, revealing vulnerabilities in guardrails and identifying strategies like gradual persuasion, splitting conversations, pre-prompting, and manipulating settings. Experimental testing of ChatGPT and Gemini demonstrated how these tools could enable harassment, surveillance, and gaslighting. Recommendations include analyzing multi-session user patterns to detect coercive control and ensuring visibility of pre-programmed settings to prevent covert manipulation. These insights aim to make AI tools safer and less susceptible to abuse.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 63%
Dark Patterns after the GDPR: Scraping Consent Pop-ups and Demonstrating their Influence
CHI '20· AI Ethics, Fairness & Accountability +2
- 63%
Is this AI trained on Credible Data? The Effects of Labeling Quality and Performance Bias on User Trust
CHI '23· Explainable AI (XAI) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)