"Are You Really Sure?'' Understanding the Effects of Human Self-Confidence Calibration in AI-Assisted Decision Making
Authors
Document Title
“Are You Really Sure?” Understanding the Effects of Human Self-Confidence Calibration in AI-Assisted Decision Making
Document Information
- Subject Area: Human-Computer Interaction, AI Decision Support Systems
- Keywords: AI-assisted decision-making, human-AI collaboration, reliance calibration, trust calibration, appropriate reliance
Research Background and Problem
- Issues or Challenges Identified by the Authors:
- Ensuring appropriate human reliance on AI in AI-assisted decision-making is critical but challenging.
- Existing research primarily focuses on the presentation of AI confidence while neglecting the role of human self-confidence in decision-making.
- Humans often exhibit miscalibrated confidence (overconfidence or underconfidence), which affects appropriate reliance on AI.
- Why This Problem Is Important:
- Over-reliance on AI can lead to erroneous decisions, while under-reliance may result in missed correct recommendations.
- Improving the overall performance of human-AI teams requires enhancing the accuracy of human confidence.
- Research Motivation:
- By calibrating human self-confidence, the study aims to improve the efficiency and rationality of AI-human collaborative decision-making.
- To explore how different calibration mechanisms influence human confidence, reliance on AI, and task performance.
Solution
Methods and Framework
-
Analytical Framework:
- The study proposes the “Confidence-Correctness Matching” (C-C Matching) framework to analyze whether the confidence levels of both humans and AI align with their actual accuracy, thereby explaining the causes of inappropriate reliance in decision-making.
-
Research Questions:
- RQ1: How does mismatched human confidence affect reliance on AI recommendations?
- RQ2: How can human confidence be calibrated, and how do calibration mechanisms impact user experience?
- RQ3: How do calibration mechanisms influence decision performance and the appropriateness of reliance on AI?
-
Calibration Mechanisms:
- Think the Opposite: Encourages users to deeply consider the scenario where their current prediction is incorrect.
- Thinking in Bets: Uses a “betting” mechanism to improve the accuracy of human confidence assessments.
- Calibration Status Feedback: Provides real-time or cumulative feedback on whether the user’s confidence matches their accuracy.
-
Experimental Design:
- Three studies were conducted: (1) Analyzing the relationship between confidence and decision reliance; (2) Comparing the effects of three calibration mechanisms; (3) Exploring the role of calibration in human-AI collaboration.
- The experiments used an income prediction task to test the models and their effectiveness (based on a publicly available adult income dataset).
Research Findings
Specific Results
-
Study 1 (RQ1: Understanding the Importance of Confidence Calibration):
- Mismatched human confidence significantly increases decision error rates.
- Displaying AI confidence does not significantly improve human confidence calibration or decision performance.
-
Study 2 (RQ2: Comparing Different Calibration Mechanisms):
- “Think the Opposite” and “Calibration Feedback” significantly reduce confidence mismatch (measured by ECE).
- “Betting” has a slight effect on mitigating overconfidence but does not effectively improve confidence calibration.
- In terms of user experience, “Think the Opposite” increases cognitive load and reduces user satisfaction.
-
Study 3 (RQ3: The Impact of Calibration on AI-Assisted Decision-Making):
- Calibration mechanisms reduce user under-reliance but do not mitigate over-reliance.
- Calibration significantly improves initial decision accuracy and the overall performance of human-AI teams.
- When AI confidence is accurate, calibration mechanisms reduce error rates; however, when AI confidence is inaccurate, calibration may increase error rates.
Advantages Over Existing Methods
- Introduces a human-centered perspective on “human confidence calibration,” addressing gaps in prior research focused on AI confidence.
- The proposed analytical framework identifies the root causes of inappropriate reliance, providing insights for improving future AI decision-making designs.
Limitations and Future Directions
- Current calibration mechanisms have limitations (e.g., high cognitive load, dependence on data feedback). Future research could explore more streamlined and efficient calibration approaches.
- The effectiveness of calibration mechanisms is highly influenced by the accuracy of AI confidence. Additional mechanisms are needed when AI confidence is unreliable.
- Calibration methods are more suitable for multi-step decision-making scenarios. Future studies could extend the approach to low-interaction, real-time decision contexts.
Conclusion
This paper reexamines the issue of reliance in AI-assisted decision-making from the perspective of user confidence, proposing a general analytical framework and experimentally validating the effectiveness and limitations of human confidence calibration. The findings offer significant theoretical and practical guidance for designing effective AI interaction interfaces.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How does human confidence miscalibration affect reliance on AI recommendations?Category: Confidence Expression and Metacognitive CalibrationSimilar questionsarrow_forward
- How can human confidence be calibrated, and what impact do calibration mechanisms have on UX?Category: Confidence Expression and Metacognitive CalibrationSimilar questionsarrow_forward
- How do calibration mechanisms affect decision performance and appropriateness of AI reliance?Category: Confidence Expression and Metacognitive CalibrationSimilar questionsarrow_forward
Practical Problems
1- Users struggle to rely on AI appropriately; over- or under-reliance leads to decision errors.Category: Confidence Expression and Metacognitive CalibrationSimilar questionsarrow_forward
- 83%
Evaluating the Interpretability of Generative Models by Interactive Reconstruction
CHI '21· Explainable AI (XAI) +2
- 83%
PaTAT: Human-AI Collaborative Qualitative Coding with Explainable Interactive Rule Synthesis
CHI '23· Explainable AI (XAI) +2
- 83%
Faulty or Ready? Handling Failures in Deep-Learning Computer Vision Models until Deployment: A Study of Practices, Challenges, and Needs
CHI '23· Explainable AI (XAI) +2
- 83%
"Should I Rely on You or the AI?" Leaders' Trust and Perceptions in Mixed Human-AI Teams
CHI '26· Human-Robot Collaboration (HRC) +2
- 83%
Do People Appropriately Rely on AI-Advice? An Analytical Review of HCI Research on Human-AI Decision-Making
CHI '26· AI-Assisted Decision-Making & Automation +2
- 83%
Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving Tasks
CHI '26· Human-LLM Collaboration +2
- 83%
The Impact of Response Latency and Task Type on Human-LLM Interaction and Perception
CHI '26· Human-LLM Collaboration +2
- 83%
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
CHI '26· Human-LLM Collaboration +2
- 83%
Guidance Source Matters: How Guidance from AI, Expert, or a Group of Analysts Impacts Visual Data Preparation and Analysis
IUI '25· Generative AI (Text, Image, Music, Video) +2
- 83%
Unakite: Scaffolding Developers’ Decision-Making Using the Web
UIST '19· Explainable AI (XAI) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)