Emulating Aggregate Human Choice Behavior and Biases with GPT Conversational Agents
Honorable MentionPaper Title
Emulating Aggregate Human Choice Behavior and Biases with GPT Conversational Agents
Publication Info
- Topic area: Cognitive biases in decision-making and their emulation using large language models (LLMs).
- Keywords: Cognitive biases, status quo bias, large language models, GPT-4, GPT-5, conversational agents, decision-making, human behavior modeling, cognitive load, human-likeness prompts.
Background and Problem
- Problem / challenge: Previous studies have shown that LLMs can reproduce cognitive biases in isolated settings, but their ability to emulate human decision-making biases in conversational contexts influenced by cognitive load remains unexplored.
- Significance: Understanding how LLMs emulate human biases is crucial for designing adaptive, bias-aware AI systems that interact effectively in real-world scenarios such as e-commerce, healthcare, and policy-making.
- Motivation and related work: Prior research has demonstrated LLMs' ability to replicate human-like dialogue and cognitive biases in single-shot or non-interactive settings. However, these studies lack investigation into how biases interact with contextual factors like dialogue complexity and cognitive load. This paper fills this gap by examining whether LLMs can emulate human decision-making biases probabilistically in conversational contexts.
Solution
- Proposed approach: The study investigates whether LLMs (GPT-4 and GPT-5) can emulate human decision-making biases, particularly the status quo bias, in conversational settings influenced by cognitive load. It uses human baseline experiments and LLM simulations with varying levels of human-likeness prompts.
- Novelty:
- Empirical investigation of LLMs' ability to emulate human decision-making biases in conversational contexts.
- Analysis of how cognitive load interacts with biases in human and LLM responses.
- Introduction of human-likeness prompts to test LLMs' alignment with human behavior.
- Comparison of LLM performance across models and prompts to evaluate predictive accuracy and bias reproduction.
- Procedure and key techniques:
- Conduct human experiments (N=1100) using chatbot-mediated decision scenarios under simple and complex dialogue conditions to establish a baseline for status quo bias.
- Simulate human participants using GPT-4 and GPT-5 agents, leveraging demographic data and prior dialogue transcripts.
- Evaluate LLMs' predictive accuracy and ability to reproduce human biases at individual and sample levels across three human-likeness prompt levels.
- Perform ablation and perturbation studies to identify key components influencing LLM behavior.
Results
- Concrete findings:
- Human participants exhibited strong status quo bias in decision scenarios, with selection rates increasing significantly when options were framed as the status quo (e.g., 51% vs. 14% in Budget Allocation scenario, p < 0.001).
- LLMs reproduced human biases at the sample level with high precision under neutral prompts (HL1 and HL2), but overestimated biases under explicit bias-susceptibility prompts (HL3).
- Cognitive load induced by complex dialogues significantly increased perceived mental demand and effort (NASA-TLX scores: Mental Demand = 3.28 vs. 1.97, p < 0.001).
- LLMs showed limited ability to emulate individual-specific decision tendencies, relying more on general patterns in dialogue tasks and decision scenarios.
- Advantage over baselines:
- GPT-4.1 models demonstrated higher precision (up to 68.5%) and better alignment with human biases compared to GPT-5 models.
- Human-likeness prompts (HL1 and HL2) improved sample-level bias reproduction without overfitting, unlike HL3.
- Experiments / evaluation:
- Human experiments tested status quo bias across three decision scenarios (Budget Allocation, Investment Portfolios, College Jobs) under simple and complex dialogue conditions.
- LLM simulations replicated the experimental setup with varying human-likeness prompts and analyzed predictive accuracy, bias reproduction, and contextual interactions.
- Statistical models (GLMM) and metrics like precision, recall, and F1 scores were used to evaluate alignment between human and LLM responses.
- Limitations and future work:
- Findings may not generalize to real-world, domain-specific conversations due to the abstract nature of decision scenarios.
- Study focused solely on status quo bias; other biases like anchoring and framing require further investigation.
- Cognitive load was assessed using self-reported and behavioral indicators; future studies could incorporate physiological measures for real-time assessment.
- Experiments were limited to GPT-4 and GPT-5 models; broader benchmarking across other LLMs is needed.
Summary
This study demonstrates that LLMs can emulate human decision-making biases, particularly the status quo bias, in conversational contexts influenced by cognitive load. Human experiments established a baseline for bias, while LLM simulations showed high precision in reproducing sample-level biases under neutral prompts. However, LLMs struggled to capture individual-specific tendencies and overestimated biases under explicit bias-susceptibility prompts. Findings highlight the potential of LLMs for realistic behavioral simulations and adaptive, bias-aware systems, but further research is needed to generalize results to other biases and real-world applications.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 86%
Knowing About Knowing: An Illusion of Human Competence Can Hinder Appropriate Reliance on AI Systems
CHI '23· Explainable AI (XAI) +2
- 86%
Are Two Heads Better Than One in AI-Assisted Decision Making? Comparing the Behavior and Performance of Groups and Individuals in Human-AI Collaborative Recidivism Risk Assessment
CHI '23· Human-LLM Collaboration +2
- 86%
What is Human-Centered about Human-Centered AI? A Map of the Research Landscape
CHI '23· Human-LLM Collaboration +2
- 86%
Understanding Compliance and Conversion Dynamics in Multi-Agent Collectives
CHI '26· Human-LLM Collaboration +2
- 86%
A Survey of Collaborative Reinforcement Learning: Interactive Methods and Design Patterns
DIS '21· Human-LLM Collaboration +2
- 86%
Who Needs What Explanation? How User Traits Affect Explanation Effectiveness in AI-Assisted Decision-Making
IUI '26· AI-Assisted Decision-Making & Automation +2
- 75%
Explanations, Fairness, and Appropriate Reliance in Human-AI Decision-Making
CHI '24· Explainable AI (XAI) +2
- 71%
"Why is 'Chicago' deceptive?" Towards Building Model-Driven Tutorials for Humans
CHI '20· Human-LLM Collaboration +2
- 71%
You Complete Me: Human-AI Teams and Complementary Expertise
CHI '22· Human-LLM Collaboration +1
- 71%
Farsight: Fostering Responsible AI Awareness During AI Application Prototyping
CHI '24· Explainable AI (XAI) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)