Who Does What? Archetypes of Roles Assigned to LLMs During Human-AI Decision-Making
Authors
Paper Title
Who Does What? Archetypes of Roles Assigned to LLMs During Human-AI Decision-Making
Publication Info
- Topic area: Socio-technical patterns in human-LLM decision-making interactions.
- Keywords: Large language models, human-in-the-loop, decision-making, archetypes, socio-technical systems, explainability, cognitive forcing, group dynamics, AI roles, human-AI collaboration.
Background and Problem
- Problem / challenge: Current evaluations of LLMs focus on static benchmarks and isolated capabilities, overlooking the socio-technical dynamics of human-LLM interactions in real-world decision-making workflows.
- Significance: Understanding these dynamics is critical for designing responsible and effective human-LLM systems, particularly in high-stakes domains like medicine and finance.
- Motivation and related work: Existing research highlights issues like LLM hallucination, overreliance, and biases but lacks a comprehensive framework to analyze how LLMs are integrated into decision-making processes. This paper builds on prior work in explainability, prompt engineering, and socio-technical systems to address these gaps.
Solution
- Proposed approach: Introduction of a framework for "human-LLM archetypes"—recurring socio-technical patterns that define roles and interactions between humans and LLMs in decision-making.
- Novelty:
- Development of a taxonomy of 17 distinct human-LLM archetypes through thematic analysis of 113 decision-making papers.
- Empirical evaluation of archetypes in a clinical diagnostic case study, comparing their effects on LLM outputs (e.g., accuracy, agreement, explanation quality).
- Identification of seven critical design dimensions for human-LLM systems, including decision control, cognitive processes, and social positioning.
- Procedure and key techniques:
- Conducted a scoping literature review across 113 papers to identify archetypes.
- Thematically analyzed archetypes based on roles, interaction patterns, and socio-technical factors.
- Applied archetypes to a clinical decision-making task using GPT-4o, systematically varying prompts to evaluate their impact on accuracy, agreement, and explanation quality.
- Synthesized findings into actionable design guidelines for human-LLM systems.
Results
- Concrete findings:
- Archetypes like Role Taker, Model, and Judge showed varying accuracy (e.g., Judge positive: 95%, Judge negative: 87%).
- Explanation quality varied across archetypes, with Communicator outputs being more readable (Flesch score: 61.4) than External Explainer outputs (Flesch score: 19.8).
- Agreement rates depended on archetype-specific cues, such as reference judgments and social roles (e.g., Second Opinion positive: 60% agreement).
- Advantage over baselines:
- Demonstrated that archetypes influence key decision-making dimensions (accuracy, agreement, explanation complexity) beyond static benchmark evaluations.
- Highlighted trade-offs in archetype selection, such as balancing cognitive load and decision control.
- Experiments / evaluation:
- Task: Diagnosis of pulmonary embolism from radiology reports.
- Dataset: 100 reports sampled from the INSPECT EHR dataset (50 positive, 50 negative cases).
- Metrics: Accuracy, sensitivity, specificity, ROUGE scores, readability scores.
- Model: GPT-4o queried via Azure Government endpoints.
- Limitations and future work:
- Limited to one model (GPT-4o) and a single decision-making task.
- Automated metrics for explanation quality may not fully capture human-perceived accuracy or utility.
- Future work should explore additional archetypes, domains, and qualitative user studies.
Summary
This paper introduces a taxonomy of 17 human-LLM archetypes, derived from a scoping review of 113 papers, to analyze socio-technical interaction patterns in decision-making. A clinical case study demonstrates how archetype selection affects LLM outputs, including accuracy, agreement, and explanation quality. The study identifies seven critical design dimensions, such as decision control and cognitive processes, that influence the effectiveness and appropriateness of human-LLM systems. These findings provide a foundation for designing responsible and context-sensitive human-AI decision-making frameworks.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
Exploring Customizable Interactive Tools for Therapeutic Homework Support in Mental Health Counseling
CHI '26· Human-LLM Collaboration +2
- 71%
When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being
CHI '26· Human-LLM Collaboration +2
- 71%
Digitizing the Pre-consultation Experience: Impacts and Design Recommendations
CHI '26· Human-LLM Collaboration +2
- 71%
Towards Better Health Conversations: The Benefits of Context-seeking
CHI '26· Human-LLM Collaboration +2
- 67%
How Much Decision Power Should (A)I Have?: Investigating Patients’ Preferences Towards AI Autonomy in Healthcare Decision Making
CHI '24· AI-Assisted Decision-Making & Automation +1
- 67%
Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks
IUI '25· Human-LLM Collaboration +1
- 63%
Beyond Euphemisms: Rethinking LLMs for SRH in Conservative Contexts
CHI '26· Human-LLM Collaboration +3
- 63%
Prompting, Oversight, and Adoption: Physicians’ Use of Large Language Models for Diagnostic Reasoning in an LMIC
CHI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)