Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance
Authors
Title of the Paper
Can the Whole Outperform the Parts? The Impact of AI Explanations on Complementary Team Performance
Paper Information
- Subject Area: Explainable AI (XAI), Human-AI Collaboration, and Augmented Intelligence
- Keywords: AI explanations, human-AI teams, augmented intelligence, complementary performance, trust calibration, AI system reliability
Research Background and Problem
-
Observed Issues or Challenges:
- Numerous studies suggest that decision-making efficiency in human-AI teams improves when AI systems provide explanations, but these experimental scenarios often assume AI performance far surpasses human performance.
- Few studies explore whether explanations can enhance complementary performance (i.e., team performance exceeding that of humans or AI working alone) when AI accuracy is comparable to human accuracy.
- Current AI explanations may lead to inappropriate trust, resulting in blind reliance on or complete disregard for AI.
-
Why It Matters:
- In high-stakes scenarios such as medical diagnosis or judicial recommendations, maintaining human accountability for decisions is critical, making the improvement of complementary human-AI collaboration essential.
- Existing XAI research overlooks the design of scenarios focusing on the collaborative potential between AI and humans.
- Investigating how AI explanations influence trust dynamics can inform the design of more reliable and transparent human-AI systems.
-
Motivation and Related Work:
- Existing research shows that AI explanations improve team performance in certain scenarios, but experiments often focus on conditions where AI performance far exceeds human performance, neglecting more realistic settings.
- A review of prior work indicates that when AI accuracy is comparable to human accuracy, it remains an open question whether explanations, beyond showing predictions and confidence scores, can effectively enhance decision accuracy.
Solution
-
Proposed Methods/Solutions:
- Adjust experimental scenarios so that AI and human accuracy are comparable (by intentionally selecting task samples where AI performance is similar to human performance).
- Explore the impact of different explanation strategies (e.g., single-class explanations, multi-class explanations) and sources (AI-generated or expert-generated).
- Propose an adaptive explanation method, which provides multi-class explanations when AI confidence is low and single-class explanations when confidence is high.
-
Innovations:
- Introduced a novel experimental framework specifically designed to explore complementary team performance, rather than merely evaluating the presence of explanations on human-AI team performance.
- Quantified the positive and negative impacts of AI explanations on team decision-making performance, particularly highlighting the potential harm when AI is incorrect.
- Designed a new explanation mechanism (adaptive explanation) to enhance the flexibility of trust interaction with humans.
-
Implementation Steps and Key Techniques:
- Selected three task datasets (sentiment analysis, logical reasoning problems).
- Compared multiple experimental conditions: baseline conditions (showing only predictions and confidence) versus explanation conditions (including AI- or expert-generated explanations).
- Evaluated AI performance using a RoBERTa-based text classifier and generated explanations using the LIME method.
- Conducted large-scale user experiments across different tasks, collecting team performance and user feedback via crowdsourcing.
Research Findings
-
Specific Findings:
- Across all tasks, human-AI teams demonstrated complementary performance (e.g., team decision accuracy exceeded that of humans or AI working alone).
- Limited pure gain from AI explanations: Explanations did not significantly improve team performance; simply showing predictions and confidence scores achieved similar results.
- While explanations improved accuracy when AI predictions were correct, they reduced decision performance when AI was wrong, resulting in minimal net gain.
- Adaptive explanations moderately reduced blind trust in low-confidence AI but did not significantly enhance overall team performance.
-
Comparative Advantages:
- Beyond the existing research framework that primarily evaluates whether explanations “improve performance,” this study delves deeper into when, how, and why explanations can or cannot enhance complementary team performance.
- Highlights the potential of displaying AI confidence as a simple and effective approach to human-AI collaboration.
-
Experimental or Evaluation Results:
- Comprehensive evaluations revealed that even with high-quality (expert-generated) explanations, improvements in team performance were limited.
- Qualitative analysis showed that human users often used AI predictions as a reference and developed confidence-based usage strategies, but AI errors could lead to misplaced trust.
-
Limitations and Future Directions:
- This study evaluated only three types of tasks; future research should test more real-world high-stakes scenarios (e.g., medical or legal).
- Current explanation methods are primarily static; future work could explore dynamic, interactive methods (e.g., explanations tailored to user confusion).
- Investigate how to design explanations that help humans better identify AI errors, such as providing “counter-arguments” or “evidence-weighted” explanations.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- When AI accuracy is comparable to humans, can AI explanations improve complementary performance of human-AI teams (overall better than either alone)?Category: XAI Explanation and Appropriate Reliance CalibrationSimilar questionsarrow_forward
- How do different AI explanation strategies (e.g., single-class, multi-class) affect trust calibration and decision performance?Category: XAI Explanation and Appropriate Reliance CalibrationSimilar questionsarrow_forward
- Can adaptive explanations (dynamically adjusted based on AI confidence) reduce blind trust and improve team performance?Category: XAI Explanation and Appropriate Reliance CalibrationSimilar questionsarrow_forward
Practical Problems
1- AI system explanations may cause humans to over-trust or completely distrust them, affecting collaboration effectiveness.Category: XAI Explanation and Appropriate Reliance CalibrationSimilar questionsarrow_forward
- 100%
One AI Does Not Fit All: A Cluster Analysis of the Laypeople’s Perception of AI Roles
CHI '23· Explainable AI (XAI) +1
- 100%
"Help Me Help the AI": Understanding How Explainability Can Support Human-AI Interaction
CHI '23· Explainable AI (XAI) +1
- 100%
Editable XAI: Toward Bidirectional Human-AI Alignment with Co-Editable Explanations of Interpretable Attributes
CHI '26· Explainable AI (XAI) +1
- 100%
Emergent, not Immanent: A Baradian Reading of Explainable AI
CHI '26· Explainable AI (XAI) +1
- 100%
Automated Rationale Generation: a Technique for Explainable AI and its Effects on Human Perceptions
IUI '19· Explainable AI (XAI) +1
- 100%
The Effects of Example-Based Explanations in a Machine Learning Interface
IUI '19· Explainable AI (XAI) +1
- 100%
Are Explanations Helpful? A Comparative Study of the Effects of Explanations in AI-Assisted Decision-Making
IUI '21· Explainable AI (XAI) +1
- 100%
Benefits of Machine Learning Explanations: Improved Learning in an AI-assisted Sequence Prediction Task
IUI '25· Explainable AI (XAI) +1
- 100%
Personalising AI Assistance Based on Overreliance Rate in AI-Assisted Decision Making
IUI '25· Explainable AI (XAI) +1
- 100%
Counterfactual Explanations May Not Be the Best Algorithmic Recourse Approach
IUI '25· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)