Who Should I Trust: AI or Myself? Leveraging Human and AI Correctness Likelihood to Promote Appropriate Trust in AI-Assisted Decision-Making
Authors
Title of the Paper
Who Should I Trust: AI or Myself? Leveraging Human and AI Correctness Likelihood to Promote Appropriate Trust in AI-Assisted Decision-Making
Paper Information
- Field of Study: AI-Assisted Decision-Making and Human-Computer Interaction
- Keywords: AI-Assisted Decision-Making, Human-AI Collaboration, AI Trust, Trust Calibration, Decision Support
Research Background and Problem Statement
-
Identified Problems or Challenges:
- Artificial intelligence (AI) has not yet achieved 100% accuracy in many real-world applications. Solely relying on AI decisions in high-risk domains (e.g., medicine and criminal justice) is risky.
- The goal of AI-assisted decision-making is to achieve complementary human-AI performance, where the overall accuracy of team decisions exceeds that of either humans or AI acting alone.
- Current research often uses AI confidence scores as a proxy for its correctness likelihood (CL) to calibrate human trust in AI, but it neglects the estimation of human CL, which may lead to suboptimal team decision performance.
-
Research Motivation and Related Work:
- Human CL is often expressed as subjective confidence estimates, which are frequently poorly calibrated or inaccurate.
- Most existing AI trust calibration methods focus solely on AI CL, ignoring the role of human CL. When AI CL is low but human CL is even lower, encouraging humans to doubt AI may be suboptimal.
Proposed Solution
-
Proposed Framework: The authors propose a framework that integrates both human and AI CL at the task-instance level to calibrate human trust in AI.
- They introduce a method to estimate human CL, predicting human performance on new task instances.
- They design three CL utilization strategies (Direct Display, Adaptive Workflow, Adaptive Recommendation) to calibrate user trust in AI.
-
Innovative Contributions:
- The authors are the first to model and utilize both human and AI CL at the task-instance level.
- They propose an interactive rule-creation interface to simulate user decision models, overcoming the limitations of relying solely on heuristic rules.
- They introduce an "adaptive" workflow in the experimental design to dynamically alter decision presentation, enhancing the precision of human trust calibration.
-
Implementation Steps and Key Techniques:
- Human CL Modeling: Calculate human CL using performance on similar tasks, and generate decision rule models through data-driven initialization and interactive refinement.
- CL Utilization Strategy Design:
- Direct Display: Directly display human and AI CL along with AI recommendations.
- Adaptive Workflow: Dynamically adjust workflows based on human and AI CL, e.g., requiring users to make decisions first when human CL is high.
- Adaptive Recommendation: Provide only AI explanation (not specific recommendations) when human CL is high.
Research Outcomes
-
Specific Results:
- The proposed human CL modeling method effectively captures CL and predicts complementary task instances.
- Experiments demonstrate that the three CL utilization strategies outperform traditional AI confidence score methods in promoting appropriate human trust in AI.
- Preliminary exploration indicates that participants find the interactive rule interface more intuitive than decision-tree-based interfaces.
-
Comparative Advantages Over Existing Solutions:
- The proposed method considers both AI and human capabilities during trust calibration.
- The Adaptive Workflow strategy effectively reduces over-reliance on AI and enhances overall team performance through dynamic workflows.
-
Experimental or Evaluation Results:
- In experiments, the team performance under Direct Display, Adaptive Workflow, and Adaptive Recommendation strategies significantly outperformed the AI confidence score method (improvement of approximately 3-4%).
- When AI confidence is inconsistent with prediction outcomes (e.g., high confidence but incorrect prediction), the new strategies exhibit higher accuracy and robustness.
- In terms of subjective perception, Adaptive Workflow received higher acceptance and perceived utility from participants regarding CL information.
-
Limitations and Future Directions:
- The current method relies on user-provided rules, requiring more generalizable modeling approaches for complex data (e.g., images and text).
- The current CL modeling does not account for dynamic changes in human decision models over long-term tasks.
- The method has not yet been validated in high-risk task scenarios (e.g., clinical diagnosis).
- Ethical and accountability issues remain, such as the potential for inaccurate human CL estimates leading to erroneous decisions.
Future Research
- Explore how to dynamically update user decision models to adapt to long-term human-AI interactions.
- Extend decision models to handle high-complexity tasks involving text and images.
- Investigate how interpretative designs can help different users better understand CL information, improving acceptance and effectiveness.
- Validate the feasibility and safety of the method in high-risk task environments, while exploring the visualization of uncertainty information.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- In AI-assisted decision-making, can combining human correctness likelihood (CL) and AI correctness likelihood (CL) improve overall team decision accuracy?Category: AI Decision Support and Reliance BehaviorSimilar questionsarrow_forward
- How can human correctness likelihood be effectively modeled and utilized in specific task contexts?Category: AI Decision Support and Reliance BehaviorSimilar questionsarrow_forward
- Which strategies for calibrating user trust in AI can optimize human-AI collaboration performance?Category: AI Decision Support and Reliance BehaviorSimilar questionsarrow_forward
Practical Problems
1- Users struggle to appropriately trust or question AI in AI-assisted decision-making, affecting decision accuracy.Category: AI Decision Support and Reliance BehaviorSimilar questionsarrow_forward
- 83%
Rethinking User Empowerment in AI Recommender System: Innovating Transparent and Controllable Interfaces
CHI '26· Explainable AI (XAI) +2
- 83%
Interactive Explainable Ranking
CHI '26· Explainable AI (XAI) +2
- 80%
Predicting and Explaining Mobile UI Tappability with Vision Modeling and Saliency Analysis
CHI '22· Explainable AI (XAI) +1
- 80%
The Role of Initial Acceptance Attitudes Toward AI Decisions in Algorithmic Recourse
CHI '25· Explainable AI (XAI) +1
- 80%
The Amplifying Effect of Explainability in AI-assisted Decision-making in Groups
CHI '25· Explainable AI (XAI) +1
- 80%
Underspecified Human Decision Experiments Considered Harmful
CHI '25· Explainable AI (XAI) +1
- 80%
Guided Reflection in AI-Assisted Decision-Making: Effects on AI Overreliance and Decision Accuracy
CHI '26· AI-Assisted Decision-Making & Automation +1
- 80%
Understanding the Effects of AI-Assisted Critical Thinking on Human-AI Decision Making
CHI '26· AI-Assisted Decision-Making & Automation +1
- 71%
Solving Separation-of-Concerns Problems in Collaborative Design of Human-AI Systems through Leaky Abstractions
CHI '22· Human-LLM Collaboration +3
- 67%
Considering Agency and Data Granularity in the Design of Visualization Tools
CHI '18· Explainable AI (XAI) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)