More Isn't Always Better: Balancing Decision Accuracy and Conformity Pressures in Multi-AI Advice
Paper Title
More Isn't Always Better: Balancing Decision Accuracy and Conformity Pressures in Multi-AI Advice
Publication Info
- Topic area: Human decision-making with multi-AI systems
- Keywords: Multi-AI advice, conformity pressure, decision accuracy, human-AI interaction, panel size, anthropomorphism, reliance, consensus, cognitive load, minority dissent
Background and Problem
- Problem / challenge: Existing research on human-AI decision-making primarily focuses on single AI advisors, leaving unclear how advice from multiple AI systems affects decision accuracy and reliance. Conformity pressure and cognitive overload are potential risks in multi-AI settings.
- Significance: Understanding how multi-AI advice influences human decisions is critical for designing effective decision-support systems in domains like medicine, law, and education, where accuracy and trust are essential.
- Motivation and related work: Prior work has shown that aggregating multiple opinions can improve accuracy but also risks conformity and confusion. Studies on human-AI teams and ensemble models have not sufficiently addressed the dynamics of multiple independent AI advisors, leaving a gap in understanding how multi-AI panels impact human decision-making.
Solution
- Proposed approach: The study investigates the effects of AI panel size, within-panel consensus, and human-likeness of AI advisors on decision accuracy, reliance, and conformity pressure.
- Novelty:
- Empirical analysis of multi-AI panel size effects on decision accuracy and reliance.
- Examination of within-panel consensus and its impact on conformity pressure and decision-making.
- Exploration of human-likeness in AI presentation and its influence on reliance and perceived autonomy.
- Procedure and key techniques:
- Conducted two studies with 348 participants across three decision tasks (Income, Recidivism, Dating).
- Manipulated panel size (1, 3, or 5 AIs), consensus levels, and human-likeness of AI advisors.
- Used random forest models for AI predictions and GPT-4o for generating natural-language explanations.
- Measured decision accuracy, reliance, confidence changes, and subjective perceptions using statistical tests and questionnaires.
Results
- Concrete findings:
- Medium panel size (three AIs) improved accuracy compared to a single AI, while larger panels (five AIs) offered no additional benefit.
- High within-panel consensus induced strong conformity pressure, improving accuracy when AIs were correct but risking overreliance. Minority dissent reduced conformity but wide disagreement caused confusion.
- Human-like AI presentations increased perceived autonomy and usefulness in certain tasks but did not significantly affect accuracy or reliance overall.
- Advantage over baselines:
- Three-AI panels consistently outperformed single-AI panels in accuracy for Income and Dating tasks.
- Minority dissent mitigated overreliance compared to unanimous AI advice, promoting more balanced decision-making.
- Experiments / evaluation:
- Tasks calibrated for 60–70% human-only accuracy to ensure appropriate difficulty.
- Statistical analysis of reliance indicators (e.g., Agreement Fraction, RAIR, RSR) and subjective measures (e.g., conformity pressure, perceived usefulness).
- Manipulation check confirmed successful induction of human-likeness perceptions in Study 2.
- Limitations and future work:
- Tasks were short and text-based, limiting generalization to real-world, high-stakes decisions.
- Participants were primarily from Asia, requiring broader samples for cultural generalization.
- Future work should explore richer interfaces (e.g., voice-based systems) and behavioral methods (e.g., eye-tracking) to analyze cognitive load and conformity dynamics.
Summary
This study examined how advice from multiple AI panels affects human decision-making, focusing on panel size, consensus, and human-likeness. Results showed that medium-sized panels (three AIs) improved accuracy, while larger panels offered no additional benefit. High consensus induced conformity pressure, improving accuracy when AIs were correct but risking overreliance. Minority dissent mitigated conformity and promoted balanced reliance, while wide disagreement caused confusion. Human-likeness enhanced perceived autonomy and usefulness in specific contexts but did not significantly affect accuracy. These findings highlight the need for calibrated panel sizes, structured presentation of consensus, and context-sensitive anthropomorphic designs to optimize multi-AI decision-support systems.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
Belief Updating and Delegation in Multi-Task Human–AI Interaction: Evidence from Controlled Simulations
CHI '26· Human-LLM Collaboration +2
- 100%
Strategic Tradeoffs Between Humans and AI in Multi-Agent Bargaining
IUI '26· Human-LLM Collaboration +2
- 100%
Making Absence Visible in Intelligent Summarization Interfaces
IUI '26· Human-LLM Collaboration +2
- 100%
"Un-default" Behavior Tuning: Specifying Model Behavior outside the Norm with LLM Self-Playing and Self-Improving
IUI '26· Human-LLM Collaboration +2
- 83%
Which Contributions Deserve Credit? Perceptions of Attribution in Human-AI Co-Creation
CHI '25· Human-LLM Collaboration +2
- 83%
Co-Disclosing the Computer: LLM-Mediated Computing through Reflective Conversation
CHI '26· Human-LLM Collaboration +2
- 83%
What can AI do for me: Evaluating Machine Learning Interpretations in Cooperative Play
IUI '19· Human-LLM Collaboration +2
- 83%
CAIM: Development and Evaluation of a Cognitive AI Memory Framework for Long-Term Interaction with Intelligent Agents
IUI '26· Human-LLM Collaboration +2
- 83%
User Reliance on AI Support for Collaborative Partner Selection
IUI '26· Human-LLM Collaboration +2
- 83%
EvalAssist: Insights on Task-Specific Evaluations and AI-Assisted Judgment Strategy Preferences
UIST '25· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)