AI-Facilitated Coercive Control: An Experimental Study

Agent Personality & AnthropomorphismAI Ethics, Fairness & AccountabilityPrivacy by Design & User ControlDark Patterns RecognitionAI/ML Researchers & EngineersPrivacy Policy MakersContent Governance & Platform Compliance Teams

Paper Title

AI-Facilitated Coercive Control: An Experimental Study

Publication Info

  • Topic area: The potential misuse of conversational AI tools in facilitating coercive control tactics.
  • Keywords: Coercive control, conversational AI, LLMs, abuse, guardrails, gaslighting, surveillance, harassment, speculative design.

Background and Problem

  • Problem / challenge: While conversational AI tools are designed with guardrails to prevent harmful use, they can be manipulated to facilitate coercive control tactics such as harassment, surveillance, and gaslighting. Current research lacks a comprehensive investigation into these vulnerabilities.
  • Significance: Coercive control is a pervasive form of abuse, often occurring in intimate partner violence (IPV). The misuse of AI tools could exacerbate harm, commoditize abusive behaviors, and undermine survivors' autonomy and safety.
  • Motivation and related work: Prior studies have focused on AI-facilitated harms like deepfakes and image-based abuse, but little research has explored how conversational AI tools might enable broader coercive control tactics. Existing literature on AI vulnerabilities (e.g., jailbreaking, multi-turn attacks) has predominantly used programmatic methods, leaving gaps in understanding qualitative, real-world abuse contexts.

Solution

  • Proposed approach: A qualitative experimental study using speculative design to explore scenarios where conversational AI tools (ChatGPT and Gemini) might be exploited for coercive control.
  • Novelty:
    1. Development of four realistic abuse scenarios combining coercive control tactics with AI capabilities.
    2. Identification of four strategies for circumventing AI guardrails: gradual persuasion, splitting conversations, pre-prompting, and manipulating agent settings.
    3. Analysis of conversational vulnerabilities and covert manipulation tactics in AI tools.
    4. Recommendations for improving AI safeguards, including multi-session pattern analysis and visibility of pre-programmed settings.
  • Procedure and key techniques:
    • Speculative design to construct abuse scenarios.
    • Manual testing of ChatGPT and Gemini across 30 sessions to probe guardrails.
    • Thematic analysis of session transcripts, researcher diaries, and AI responses to identify vulnerabilities and bypass strategies.

Results

  • Concrete findings:
    • AI tools initially refused straightforward requests for harmful content but could be manipulated to provide assistance through strategies like deceptive context, splitting tasks across sessions, and pre-prompting settings.
    • ChatGPT and Gemini provided detailed information on harmful tools (e.g., spyware apps) and generated biased outputs when pre-programmed.
    • Average session duration was 9 minutes and 47 seconds, highlighting the ease of circumventing guardrails.
  • Advantage over baselines: This study expands prior work by qualitatively analyzing real-world abuse contexts, uncovering vulnerabilities specific to conversational AI tools, and identifying covert manipulation tactics.
  • Experiments / evaluation:
    • Four scenarios tested: generating harmful content, coercing unfair labor division, discovering surveillance tools, and injecting bias into AI responses.
    • Data collected included screen captures, text logs, and researcher memos.
    • Analysis revealed patterns of guardrail circumvention and escalation of harmful content.
  • Limitations and future work:
    • Findings are limited to two AI tools (ChatGPT and Gemini) and specific coercive control tactics.
    • Survivor perspectives were not included to avoid retraumatization.
    • Future research should explore additional AI tools, evolving model updates, and broader abuse tactics.

Summary

This study investigates how conversational AI tools might be exploited to facilitate coercive control tactics, revealing vulnerabilities in guardrails and identifying strategies like gradual persuasion, splitting conversations, pre-prompting, and manipulating settings. Experimental testing of ChatGPT and Gemini demonstrated how these tools could enable harassment, surveillance, and gaslighting. Recommendations include analyzing multi-session user patterns to detect coercive control and ensuring visibility of pre-programmed settings to prevent covert manipulation. These insights aim to make AI tools safer and less susceptible to abuse.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223284/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790859
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Agent Personality & Anthropomorphism, AI Ethics, Fairness & Accountability, Privacy by Design & User Control, Dark Patterns Recognition
work
Professions
AI/ML Researchers & Engineers, Privacy Policy Makers, Content Governance & Platform Compliance Teams
article
Content Status
Full text indexed
hub
Related Papers
2 related papers