CoBRA: Programming Cognitive Bias in Social Agents Using Classic Social Science Experiments

Best Paper
Human-LLM CollaborationExplainable AI (XAI)Brain-Computer Interface (BCI) & NeurofeedbackAffective Human-Computer DialogueAI/ML Researchers & EngineersHCI ResearchersCognitive Scientists

Paper Title

CoBRA: Programming Cognitive Bias in Social Agents Using Classic Social Science Experiments

Publication Info

  • Topic area: Cognitive bias control in LLM-based social simulations
  • Keywords: Cognitive bias, social agents, LLMs, behavioral regulation, reproducibility, framing effect, authority effect, bandwagon effect, confirmation bias, social science experiments

Background and Problem

  • Problem / challenge: Existing methods for specifying agent behavior in social simulations rely on implicit natural language descriptions, leading to inconsistent and unpredictable behaviors across models. These approaches fail to reliably capture nuanced, role-based differences in behavior.
  • Significance: Reliable control over agent behavior is critical for reproducible social simulations, which can be used to test social theories and explore complex human behaviors in a controlled, ethical manner.
  • Motivation and related work: Prior work has explored LLM-based agent modeling and alignment but lacks mechanisms to systematically control cognitive biases. This paper addresses the gap by introducing a toolkit that operationalizes validated social science experiments to regulate agent behavior.

Solution

  • Proposed approach: CoBRA (Cognitive Bias Regulator for Social Agents), a toolkit for explicitly and quantitatively controlling cognitive biases in LLM-based social agents.
  • Novelty:
    1. Introduces a closed-loop system for measuring and regulating cognitive biases using validated social science experiments.
    2. Develops a Cognitive Bias Index (CBI) as a quantitative metric for bias control.
    3. Implements a Behavioral Regulation Engine with three intervention methods: prompt engineering, representation engineering, and fine-tuning.
    4. Demonstrates reproducibility and controllability across models, temperatures, and reasoning modes.
  • Procedure and key techniques:
    1. Cognitive Bias Index (CBI): Measures bias levels using structured prompts derived from classic experiments (e.g., Asian Disease, Milgram Obedience).
    2. Behavioral Regulation Engine: Aligns agent behavior using:
      • Prompt Numerical Control: Directly specifies bias levels via numerical instructions.
      • Representation Engineering: Modifies activation-space representations to inject bias signals.
      • Fine-tuning: Adjusts model parameters using lightweight LoRA-based methods.
    3. Evaluates agents across multiple paradigms to ensure reproducibility and controllability.

Results

  • Concrete findings:
    • CoBRA reduces cross-model variance by 77% compared to baselines.
    • Control coefficients transfer reliably across paradigms (Spearman’s ρ > 0.96).
    • Representation Engineering achieves the best balance of monotonicity (NDCG ≈ 1.00), smoothness (Δ1 = 0.07), and expressiveness (range ≈ 1.8–2.1).
  • Advantage over baselines:
    • Significantly outperforms implicit natural language specifications in consistency and predictability.
    • Achieves stable control across temperatures (variance < 1.2%) and reasoning modes (equivalence in 100% of tests for RepE methods).
  • Experiments / evaluation:
    • Benchmarked across four cognitive biases (Authority Effect, Bandwagon Effect, Confirmation Bias, Framing Effect) using eight classic paradigms.
    • Tested on both open-source and closed-source models, with robust performance in API-only settings.
    • Demonstrated in an emotional contagion simulation, showing predictable dose–response relationships.
  • Limitations and future work:
    • Usability studies with sociologists are pending.
    • Limited to language-based simulations; future work will explore multi-modal settings.
    • Focused on single-bias control; compositional bias control remains unexplored.
    • Activation and parameter-level controls are restricted to open-source models.

Summary

CoBRA introduces a novel framework for explicitly controlling cognitive biases in LLM-based social agents using validated social science experiments. By leveraging a Cognitive Bias Index and a Behavioral Regulation Engine, CoBRA ensures reproducible and fine-grained behavior control across models, temperatures, and reasoning modes. The toolkit demonstrates significant improvements over baseline methods in consistency, predictability, and generalizability. While currently limited to language-based and single-bias simulations, CoBRA lays the groundwork for more precise and impactful social science simulations.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223158/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790804
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
Best Paper
group
Authors
3 authors
sell
Subtopics
Human-LLM Collaboration, Explainable AI (XAI), Brain-Computer Interface (BCI) & Neurofeedback, Affective Human-Computer Dialogue
work
Professions
AI/ML Researchers & Engineers, HCI Researchers, Cognitive Scientists
article
Content Status
Full text indexed
hub
Related Papers
5 related papers