When Confidence Meets Accuracy: Exploring the Effects of Multiple Performance Indicators on Trust in Machine Learning Models

Best Paper
Explainable AI (XAI)AI-Assisted Decision-Making & AutomationAI/ML Researchers & EngineersCognitive Scientists

Document Title

When Confidence Meets Accuracy: Exploring the Effects of Multiple Performance Indicators on Trust in Machine Learning Models

Document Information

  • Subject Area: Artificial Intelligence and Human-Computer Interaction
  • Keywords: Machine Learning, Confidence, Accuracy, Trust, Human Experiment, Performance Indicators

Research Background and Problem

  • Identified Problem or Challenge: With the widespread application of machine learning models in daily life and decision-making scenarios, understanding how users perceive and trust these models has become a critical issue. Although performance indicators (e.g., accuracy) have been shown to significantly influence user trust, the impact of presenting multiple performance indicators simultaneously remains unclear.
  • Importance of the Problem: Users' trust in machine learning models directly affects their adoption of the model's predictions. This issue is particularly important because different indicators may lead to inconsistent trust levels, thereby affecting the practical effectiveness of the models.
  • Research Motivation and Related Work: Existing studies suggest that a single performance indicator (e.g., accuracy or model confidence) can influence user trust. However, research on the combined effects of multiple indicators (e.g., accuracy and confidence) on user trust is still in its exploratory phase.

Solution

  • Proposed Method or Solution:
    • A randomized behavioral experiment was designed to investigate the combined effects of multiple performance indicators on user trust.
    • The experiment systematically controlled and compared factors influencing user trust, including model confidence, stated accuracy, and observed accuracy.
  • Innovative Aspects:
    • Simultaneously examined the interaction effects of multiple performance indicators on user trust.
    • Compared user responses to individual indicators versus combined indicators.
  • Implementation Steps and Key Techniques:
    • Recruited general participants via Amazon Mechanical Turk (MTurk) to complete multiple rounds of low-risk decision-making tasks, collaborating with a machine learning model to predict speed dating outcomes.
    • Tested eight experimental conditions combining different levels of confidence, stated accuracy, and observed accuracy.
    • Measured trust across multiple dimensions, including self-reported trust, adoption of model predictions, and prediction conversion rates.

Research Findings

  • Specific Findings:
    • Before users observed the model's actual performance, high-confidence models led users to believe in the accuracy of predictions. However, stated accuracy had a more significant impact on trust.
    • After users observed the model's actual performance, observed accuracy became the dominant factor influencing trust, surpassing both stated accuracy and confidence levels.
    • There was an interaction effect between model confidence and accuracy, particularly in moderating users' beliefs about model accuracy.
  • Comparison with Existing Solutions and Advantages:
    • This study not only focused on the impact of individual indicators on trust but also explored the effects of interactions between multiple indicators, addressing the limitations of prior single-indicator studies.
    • It was the first to clearly identify that observed accuracy in practice has the most significant impact on trust and may lead users to disregard model confidence.
  • Experimental or Evaluation Results:
    • Model confidence primarily influenced users' beliefs about accuracy but had minimal effect on self-reported trust levels or the frequency of adopting model predictions.
    • Among the three indicators, observed accuracy had the greatest impact on user behavior and trust.
  • Limitations and Future Directions:
    • The study was limited to extreme experimental conditions (e.g., highly differentiated accuracy and confidence levels), which may not fully reflect user responses in intermediate scenarios.
    • Findings derived from a single task type (speed dating prediction) may be difficult to generalize to high-risk decision-making contexts (e.g., medical diagnosis).
    • Future research could explore other performance indicators (e.g., F1 score) and combinations of indicators, as well as develop theoretical models to explain how people perceive, process, and respond to machine learning performance information.

Conclusion

This study provides a preliminary empirical analysis of the effects of multiple performance indicators on user trust, clarifying the scope and interaction of different indicators. The findings offer important insights for task design, result analysis, and the practical application of machine learning models. Future research should validate and extend these findings across broader application scenarios and user groups.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/68856/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/abs/10.1145/3491102.3501967
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2022
emoji_events
Award
Best Paper
group
Authors
2 authors
sell
Subtopics
Explainable AI (XAI), AI-Assisted Decision-Making & Automation
work
Professions
AI/ML Researchers & Engineers, Cognitive Scientists
article
Content Status
Full text indexed
hub
Related Papers
10 related papers