Understanding the Effect of Counterfactual Explanations on Trust and Reliance on AI for Human-AI Collaborative Clinical Decision Making
Artificial intelligence (AI) is increasingly being considered to assist human decision-making in high-stake domains (e.g. health). In particular, AI models aim to bring complementary, data-driven insights to human experts’ knowledge to improve a decision outcome that either an AI model or a human cannot achieve alone. However, researchers have discussed an issue that humans over-rely on imperfect suggestions of the AI model. In this work, we utilized salient features and what-if, counterfactual explanations to make humans review AI suggestions more thoroughly and explore the effect of these explanations on trust and reliance on AI during clinical decision-making. We conducted an experiment with seven therapists and ten novices on the task of assessing post-stroke survivors' quality of motion, in which we analyzed their performance, agreement level on the task without and with two types of AI explanations and their reliance on AI. Our results showed that the human + AI team with both salient features and counterfactual explanations assisted to improve the performance and agreement level on the task when `right' AI outputs are presented. Otherwise, both therapists and novices over-relied on `wrong' AI outputs. \red{Specifically, novices had higher performance degrades by 18.0 f1-score with salient features and 14.0 f1-score with counterfactuals than therapists with performance degrades of 8.6 and 2.8 f1-scores respectively}. When `wrong' AI outputs are presented, counterfactual explanations assisted both therapists and novices to reduce their over-reliance on `wrong' AI outputs by 21\% compared to salient feature explanations. Our work discusses the potential of counterfactual explanations to better estimate the accuracy of an AI model and reduce over-reliance on `wrong' AI outputs and implications for improving human-AI collaborative decision-making.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)