Who Should I Trust: AI or Myself? Leveraging Human and AI Correctness Likelihood to Promote Appropriate Trust in AI-Assisted Decision-Making

Explainable AI (XAI)AI-Assisted Decision-Making & AutomationUI/UX DesignersData Scientists & AnalystsAI/ML Researchers & Engineers

Title of the Paper

Who Should I Trust: AI or Myself? Leveraging Human and AI Correctness Likelihood to Promote Appropriate Trust in AI-Assisted Decision-Making

Paper Information

  • Field of Study: AI-Assisted Decision-Making and Human-Computer Interaction
  • Keywords: AI-Assisted Decision-Making, Human-AI Collaboration, AI Trust, Trust Calibration, Decision Support

Research Background and Problem Statement

  • Identified Problems or Challenges:

    1. Artificial intelligence (AI) has not yet achieved 100% accuracy in many real-world applications. Solely relying on AI decisions in high-risk domains (e.g., medicine and criminal justice) is risky.
    2. The goal of AI-assisted decision-making is to achieve complementary human-AI performance, where the overall accuracy of team decisions exceeds that of either humans or AI acting alone.
    3. Current research often uses AI confidence scores as a proxy for its correctness likelihood (CL) to calibrate human trust in AI, but it neglects the estimation of human CL, which may lead to suboptimal team decision performance.
  • Research Motivation and Related Work:

    • Human CL is often expressed as subjective confidence estimates, which are frequently poorly calibrated or inaccurate.
    • Most existing AI trust calibration methods focus solely on AI CL, ignoring the role of human CL. When AI CL is low but human CL is even lower, encouraging humans to doubt AI may be suboptimal.

Proposed Solution

  • Proposed Framework: The authors propose a framework that integrates both human and AI CL at the task-instance level to calibrate human trust in AI.

    • They introduce a method to estimate human CL, predicting human performance on new task instances.
    • They design three CL utilization strategies (Direct Display, Adaptive Workflow, Adaptive Recommendation) to calibrate user trust in AI.
  • Innovative Contributions:

    • The authors are the first to model and utilize both human and AI CL at the task-instance level.
    • They propose an interactive rule-creation interface to simulate user decision models, overcoming the limitations of relying solely on heuristic rules.
    • They introduce an "adaptive" workflow in the experimental design to dynamically alter decision presentation, enhancing the precision of human trust calibration.
  • Implementation Steps and Key Techniques:

    1. Human CL Modeling: Calculate human CL using performance on similar tasks, and generate decision rule models through data-driven initialization and interactive refinement.
    2. CL Utilization Strategy Design:
      • Direct Display: Directly display human and AI CL along with AI recommendations.
      • Adaptive Workflow: Dynamically adjust workflows based on human and AI CL, e.g., requiring users to make decisions first when human CL is high.
      • Adaptive Recommendation: Provide only AI explanation (not specific recommendations) when human CL is high.

Research Outcomes

  • Specific Results:

    1. The proposed human CL modeling method effectively captures CL and predicts complementary task instances.
    2. Experiments demonstrate that the three CL utilization strategies outperform traditional AI confidence score methods in promoting appropriate human trust in AI.
    3. Preliminary exploration indicates that participants find the interactive rule interface more intuitive than decision-tree-based interfaces.
  • Comparative Advantages Over Existing Solutions:

    1. The proposed method considers both AI and human capabilities during trust calibration.
    2. The Adaptive Workflow strategy effectively reduces over-reliance on AI and enhances overall team performance through dynamic workflows.
  • Experimental or Evaluation Results:

    • In experiments, the team performance under Direct Display, Adaptive Workflow, and Adaptive Recommendation strategies significantly outperformed the AI confidence score method (improvement of approximately 3-4%).
    • When AI confidence is inconsistent with prediction outcomes (e.g., high confidence but incorrect prediction), the new strategies exhibit higher accuracy and robustness.
    • In terms of subjective perception, Adaptive Workflow received higher acceptance and perceived utility from participants regarding CL information.
  • Limitations and Future Directions:

    1. The current method relies on user-provided rules, requiring more generalizable modeling approaches for complex data (e.g., images and text).
    2. The current CL modeling does not account for dynamic changes in human decision models over long-term tasks.
    3. The method has not yet been validated in high-risk task scenarios (e.g., clinical diagnosis).
    4. Ethical and accountability issues remain, such as the potential for inaccurate human CL estimates leading to erroneous decisions.

Future Research

  1. Explore how to dynamically update user decision models to adapt to long-term human-AI interactions.
  2. Extend decision models to handle high-complexity tasks involving text and images.
  3. Investigate how interpretative designs can help different users better understand CL information, improving acceptance and effectiveness.
  4. Validate the feasibility and safety of the method in high-risk task environments, while exploring the visualization of uncertainty information.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/96211/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581058
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Explainable AI (XAI), AI-Assisted Decision-Making & Automation
work
Professions
UI/UX Designers, Data Scientists & Analysts, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers