When Confidence Meets Accuracy: Exploring the Effects of Multiple Performance Indicators on Trust in Machine Learning Models
Best PaperDocument Title
When Confidence Meets Accuracy: Exploring the Effects of Multiple Performance Indicators on Trust in Machine Learning Models
Document Information
- Subject Area: Artificial Intelligence and Human-Computer Interaction
- Keywords: Machine Learning, Confidence, Accuracy, Trust, Human Experiment, Performance Indicators
Research Background and Problem
- Identified Problem or Challenge: With the widespread application of machine learning models in daily life and decision-making scenarios, understanding how users perceive and trust these models has become a critical issue. Although performance indicators (e.g., accuracy) have been shown to significantly influence user trust, the impact of presenting multiple performance indicators simultaneously remains unclear.
- Importance of the Problem: Users' trust in machine learning models directly affects their adoption of the model's predictions. This issue is particularly important because different indicators may lead to inconsistent trust levels, thereby affecting the practical effectiveness of the models.
- Research Motivation and Related Work: Existing studies suggest that a single performance indicator (e.g., accuracy or model confidence) can influence user trust. However, research on the combined effects of multiple indicators (e.g., accuracy and confidence) on user trust is still in its exploratory phase.
Solution
- Proposed Method or Solution:
- A randomized behavioral experiment was designed to investigate the combined effects of multiple performance indicators on user trust.
- The experiment systematically controlled and compared factors influencing user trust, including model confidence, stated accuracy, and observed accuracy.
- Innovative Aspects:
- Simultaneously examined the interaction effects of multiple performance indicators on user trust.
- Compared user responses to individual indicators versus combined indicators.
- Implementation Steps and Key Techniques:
- Recruited general participants via Amazon Mechanical Turk (MTurk) to complete multiple rounds of low-risk decision-making tasks, collaborating with a machine learning model to predict speed dating outcomes.
- Tested eight experimental conditions combining different levels of confidence, stated accuracy, and observed accuracy.
- Measured trust across multiple dimensions, including self-reported trust, adoption of model predictions, and prediction conversion rates.
Research Findings
- Specific Findings:
- Before users observed the model's actual performance, high-confidence models led users to believe in the accuracy of predictions. However, stated accuracy had a more significant impact on trust.
- After users observed the model's actual performance, observed accuracy became the dominant factor influencing trust, surpassing both stated accuracy and confidence levels.
- There was an interaction effect between model confidence and accuracy, particularly in moderating users' beliefs about model accuracy.
- Comparison with Existing Solutions and Advantages:
- This study not only focused on the impact of individual indicators on trust but also explored the effects of interactions between multiple indicators, addressing the limitations of prior single-indicator studies.
- It was the first to clearly identify that observed accuracy in practice has the most significant impact on trust and may lead users to disregard model confidence.
- Experimental or Evaluation Results:
- Model confidence primarily influenced users' beliefs about accuracy but had minimal effect on self-reported trust levels or the frequency of adopting model predictions.
- Among the three indicators, observed accuracy had the greatest impact on user behavior and trust.
- Limitations and Future Directions:
- The study was limited to extreme experimental conditions (e.g., highly differentiated accuracy and confidence levels), which may not fully reflect user responses in intermediate scenarios.
- Findings derived from a single task type (speed dating prediction) may be difficult to generalize to high-risk decision-making contexts (e.g., medical diagnosis).
- Future research could explore other performance indicators (e.g., F1 score) and combinations of indicators, as well as develop theoretical models to explain how people perceive, process, and respond to machine learning performance information.
Conclusion
This study provides a preliminary empirical analysis of the effects of multiple performance indicators on user trust, clarifying the scope and interaction of different indicators. The findings offer important insights for task design, result analysis, and the practical application of machine learning models. Future research should validate and extend these findings across broader application scenarios and user groups.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How does simultaneously presenting multiple performance metrics (e.g., confidence and accuracy) affect users' trust in machine learning models?Category: XAI Explanation and Appropriate Reliance CalibrationSimilar questionsarrow_forward
- After observing actual performance, how do users' trust weights for confidence and accuracy change?Category: XAI Explanation and Appropriate Reliance CalibrationSimilar questionsarrow_forward
- How do interactions among performance metrics shape users' beliefs about machine learning model accuracy?Category: XAI Explanation and Appropriate Reliance CalibrationSimilar questionsarrow_forward
Practical Problems
1- Users' trust in machine learning models decreases due to metric conflicts, affecting model adoption.Category: XAI Explanation and Appropriate Reliance CalibrationSimilar questionsarrow_forward
- 80%
User Characteristics in Explainable AI: The Rabbit Hole of Personalization?
CHI '24· Explainable AI (XAI) +1
- 80%
I Can Do Better Than Your AI: Expertise and Explanations
IUI '19· Explainable AI (XAI) +1
- 75%
(Mis)Communicating with our AI Systems
CHI '25· Explainable AI (XAI)
- 67%
Knowing About Knowing: An Illusion of Human Competence Can Hinder Appropriate Reliance on AI Systems
CHI '23· Explainable AI (XAI) +2
- 67%
A Survey of Collaborative Reinforcement Learning: Interactive Methods and Design Patterns
DIS '21· Human-LLM Collaboration +2
- 67%
Optimal Explanations: A Quantitative Model of Human Error in Causal Graph Interpretation
IUI '26· Explainable AI (XAI) +2
- 67%
Who Needs What Explanation? How User Traits Affect Explanation Effectiveness in AI-Assisted Decision-Making
IUI '26· AI-Assisted Decision-Making & Automation +2
- 60%
UMLAUT: Debugging Deep Learning Programs using Program Structure and Model Behavior
CHI '21· Explainable AI (XAI) +1
- 60%
Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance
CHI '21· Explainable AI (XAI) +1
- 60%
Predicting and Explaining Mobile UI Tappability with Vision Modeling and Saliency Analysis
CHI '22· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)