More Isn't Always Better: Balancing Decision Accuracy and Conformity Pressures in Multi-AI Advice

Human-LLM CollaborationAI-Assisted Decision-Making & AutomationExplainable AI (XAI)AI/ML Researchers & EngineersHCI ResearchersData Scientists & Analysts

Paper Title

More Isn't Always Better: Balancing Decision Accuracy and Conformity Pressures in Multi-AI Advice

Publication Info

  • Topic area: Human decision-making with multi-AI systems
  • Keywords: Multi-AI advice, conformity pressure, decision accuracy, human-AI interaction, panel size, anthropomorphism, reliance, consensus, cognitive load, minority dissent

Background and Problem

  • Problem / challenge: Existing research on human-AI decision-making primarily focuses on single AI advisors, leaving unclear how advice from multiple AI systems affects decision accuracy and reliance. Conformity pressure and cognitive overload are potential risks in multi-AI settings.
  • Significance: Understanding how multi-AI advice influences human decisions is critical for designing effective decision-support systems in domains like medicine, law, and education, where accuracy and trust are essential.
  • Motivation and related work: Prior work has shown that aggregating multiple opinions can improve accuracy but also risks conformity and confusion. Studies on human-AI teams and ensemble models have not sufficiently addressed the dynamics of multiple independent AI advisors, leaving a gap in understanding how multi-AI panels impact human decision-making.

Solution

  • Proposed approach: The study investigates the effects of AI panel size, within-panel consensus, and human-likeness of AI advisors on decision accuracy, reliance, and conformity pressure.
  • Novelty:
    1. Empirical analysis of multi-AI panel size effects on decision accuracy and reliance.
    2. Examination of within-panel consensus and its impact on conformity pressure and decision-making.
    3. Exploration of human-likeness in AI presentation and its influence on reliance and perceived autonomy.
  • Procedure and key techniques:
    • Conducted two studies with 348 participants across three decision tasks (Income, Recidivism, Dating).
    • Manipulated panel size (1, 3, or 5 AIs), consensus levels, and human-likeness of AI advisors.
    • Used random forest models for AI predictions and GPT-4o for generating natural-language explanations.
    • Measured decision accuracy, reliance, confidence changes, and subjective perceptions using statistical tests and questionnaires.

Results

  • Concrete findings:
    • Medium panel size (three AIs) improved accuracy compared to a single AI, while larger panels (five AIs) offered no additional benefit.
    • High within-panel consensus induced strong conformity pressure, improving accuracy when AIs were correct but risking overreliance. Minority dissent reduced conformity but wide disagreement caused confusion.
    • Human-like AI presentations increased perceived autonomy and usefulness in certain tasks but did not significantly affect accuracy or reliance overall.
  • Advantage over baselines:
    • Three-AI panels consistently outperformed single-AI panels in accuracy for Income and Dating tasks.
    • Minority dissent mitigated overreliance compared to unanimous AI advice, promoting more balanced decision-making.
  • Experiments / evaluation:
    • Tasks calibrated for 60–70% human-only accuracy to ensure appropriate difficulty.
    • Statistical analysis of reliance indicators (e.g., Agreement Fraction, RAIR, RSR) and subjective measures (e.g., conformity pressure, perceived usefulness).
    • Manipulation check confirmed successful induction of human-likeness perceptions in Study 2.
  • Limitations and future work:
    • Tasks were short and text-based, limiting generalization to real-world, high-stakes decisions.
    • Participants were primarily from Asia, requiring broader samples for cultural generalization.
    • Future work should explore richer interfaces (e.g., voice-based systems) and behavioral methods (e.g., eye-tracking) to analyze cognitive load and conformity dynamics.

Summary

This study examined how advice from multiple AI panels affects human decision-making, focusing on panel size, consensus, and human-likeness. Results showed that medium-sized panels (three AIs) improved accuracy, while larger panels offered no additional benefit. High consensus induced conformity pressure, improving accuracy when AIs were correct but risking overreliance. Minority dissent mitigated conformity and promoted balanced reliance, while wide disagreement caused confusion. Human-likeness enhanced perceived autonomy and usefulness in specific contexts but did not significantly affect accuracy. These findings highlight the need for calibrated panel sizes, structured presentation of consensus, and context-sensitive anthropomorphic designs to optimize multi-AI decision-support systems.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/221851/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791648
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
2 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation, Explainable AI (XAI)
work
Professions
AI/ML Researchers & Engineers, HCI Researchers, Data Scientists & Analysts
article
Content Status
Full text indexed
hub
Related Papers
10 related papers