Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance

Explainable AI (XAI)AI-Assisted Decision-Making & AutomationAI/ML Researchers & EngineersHCI Researchers

Title of the Paper

Can the Whole Outperform the Parts? The Impact of AI Explanations on Complementary Team Performance

Paper Information

  • Subject Area: Explainable AI (XAI), Human-AI Collaboration, and Augmented Intelligence
  • Keywords: AI explanations, human-AI teams, augmented intelligence, complementary performance, trust calibration, AI system reliability

Research Background and Problem

  • Observed Issues or Challenges:

    1. Numerous studies suggest that decision-making efficiency in human-AI teams improves when AI systems provide explanations, but these experimental scenarios often assume AI performance far surpasses human performance.
    2. Few studies explore whether explanations can enhance complementary performance (i.e., team performance exceeding that of humans or AI working alone) when AI accuracy is comparable to human accuracy.
    3. Current AI explanations may lead to inappropriate trust, resulting in blind reliance on or complete disregard for AI.
  • Why It Matters:

    1. In high-stakes scenarios such as medical diagnosis or judicial recommendations, maintaining human accountability for decisions is critical, making the improvement of complementary human-AI collaboration essential.
    2. Existing XAI research overlooks the design of scenarios focusing on the collaborative potential between AI and humans.
    3. Investigating how AI explanations influence trust dynamics can inform the design of more reliable and transparent human-AI systems.
  • Motivation and Related Work:

    • Existing research shows that AI explanations improve team performance in certain scenarios, but experiments often focus on conditions where AI performance far exceeds human performance, neglecting more realistic settings.
    • A review of prior work indicates that when AI accuracy is comparable to human accuracy, it remains an open question whether explanations, beyond showing predictions and confidence scores, can effectively enhance decision accuracy.

Solution

  • Proposed Methods/Solutions:

    1. Adjust experimental scenarios so that AI and human accuracy are comparable (by intentionally selecting task samples where AI performance is similar to human performance).
    2. Explore the impact of different explanation strategies (e.g., single-class explanations, multi-class explanations) and sources (AI-generated or expert-generated).
    3. Propose an adaptive explanation method, which provides multi-class explanations when AI confidence is low and single-class explanations when confidence is high.
  • Innovations:

    1. Introduced a novel experimental framework specifically designed to explore complementary team performance, rather than merely evaluating the presence of explanations on human-AI team performance.
    2. Quantified the positive and negative impacts of AI explanations on team decision-making performance, particularly highlighting the potential harm when AI is incorrect.
    3. Designed a new explanation mechanism (adaptive explanation) to enhance the flexibility of trust interaction with humans.
  • Implementation Steps and Key Techniques:

    1. Selected three task datasets (sentiment analysis, logical reasoning problems).
    2. Compared multiple experimental conditions: baseline conditions (showing only predictions and confidence) versus explanation conditions (including AI- or expert-generated explanations).
    3. Evaluated AI performance using a RoBERTa-based text classifier and generated explanations using the LIME method.
    4. Conducted large-scale user experiments across different tasks, collecting team performance and user feedback via crowdsourcing.

Research Findings

  • Specific Findings:

    1. Across all tasks, human-AI teams demonstrated complementary performance (e.g., team decision accuracy exceeded that of humans or AI working alone).
    2. Limited pure gain from AI explanations: Explanations did not significantly improve team performance; simply showing predictions and confidence scores achieved similar results.
    3. While explanations improved accuracy when AI predictions were correct, they reduced decision performance when AI was wrong, resulting in minimal net gain.
    4. Adaptive explanations moderately reduced blind trust in low-confidence AI but did not significantly enhance overall team performance.
  • Comparative Advantages:

    • Beyond the existing research framework that primarily evaluates whether explanations “improve performance,” this study delves deeper into when, how, and why explanations can or cannot enhance complementary team performance.
    • Highlights the potential of displaying AI confidence as a simple and effective approach to human-AI collaboration.
  • Experimental or Evaluation Results:

    1. Comprehensive evaluations revealed that even with high-quality (expert-generated) explanations, improvements in team performance were limited.
    2. Qualitative analysis showed that human users often used AI predictions as a reference and developed confidence-based usage strategies, but AI errors could lead to misplaced trust.
  • Limitations and Future Directions:

    1. This study evaluated only three types of tasks; future research should test more real-world high-stakes scenarios (e.g., medical or legal).
    2. Current explanation methods are primarily static; future work could explore dynamic, interactive methods (e.g., explanations tailored to user confusion).
    3. Investigate how to design explanations that help humans better identify AI errors, such as providing “counter-arguments” or “evidence-weighted” explanations.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/47731/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3411764.3445717
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
8 authors
sell
Subtopics
Explainable AI (XAI), AI-Assisted Decision-Making & Automation
work
Professions
AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers