Harnessing Biomedical Literature to Calibrate Clinicians' Trust in AI Decision Support Systems
Authors
Title of the Paper
Harnessing Biomedical Literature to Calibrate Clinicians’ Trust in AI Decision Support Systems
Bibliographic Information
- Research Domain: Human-Computer Interaction (HCI), Applications of Artificial Intelligence (AI) in Healthcare
- Keywords: Clinical AI, Explainable AI (XAI), Biomedical Literature, Qualitative Methods, Medical Decision Support
Research Background and Problem Statement
-
Identified Problems or Challenges:
- Clinical Decision Support Tools (DSTs) aim to assist physicians in improving diagnostic and treatment decisions. However, since AI models are not always correct, physicians need to verify AI recommendations on a case-by-case basis.
- Existing DSTs typically build trust through explanations of AI’s internal mechanisms (Explainable AI, XAI) or performance metrics. However, these methods often prove ineffective in practice, as physicians struggle to assess the accuracy of recommendations based on such explanations.
- The "opacity" and "lack of clear evidence for credibility" of AI recommendations result in incomplete trust from physicians.
-
Significance:
- Errors in AI recommendations can lead to clinical decision-making failures, potentially impacting patient health outcomes.
- Effectively calibrating physicians' trust in AI recommendations is a critical step toward expanding AI applications in healthcare.
-
Motivation and Related Work:
- The authors were inspired by the way physicians validate colleagues' recommendations in practice, often referencing scientific literature to assess the applicability and reliability of suggestions.
- Literature retrieval and evidence collection are based on principles of Evidence-Based Medicine (EBM), but AI recommendations lack explicit evidence tailored to individual patient cases.
Proposed Solution
-
Method or Solution:
- A novel decision support tool is proposed, mimicking the trust calibration interaction model among physicians that relies on literature-based evidence. This tool leverages GPT-3 to extract evidence from biomedical literature to validate the applicability and reliability of AI recommendations.
- Instead of explaining AI’s internal mechanisms, the design focuses on providing literature evidence that supports or refutes AI recommendations.
-
Innovative Features:
- Replaces existing methods of building trust through explanations of AI mechanisms with a reference to physicians' natural interaction patterns in practice.
- Emphasizes the distinction between two types of evidence: the “robustness” and “applicability” of recommendations.
- Integrates evidence from randomized clinical trials, high-quality studies, and historical cases.
-
Implementation Steps and Key Technologies:
- Utilizes a biomedical literature retrieval tool based on the PICO framework (Population, Intervention, Comparison, Outcome).
- Employs GPT-3 to provide automated summarization capabilities, generating concise literature evidence summaries for physicians to review quickly.
- Synthesizes recommendations from evidence-based medicine and AI models, aiding physicians in making informed decisions.
Research Outcomes
-
Specific Results:
- Developed a prototype DST tool that uses biomedical literature to verify the applicability and robustness of AI recommendations, helping physicians more effectively assess the accuracy of suggestions.
- Experiments revealed that physicians particularly value the applicability of AI recommendations, with matching patient samples from literature to actual patient cases being critical.
-
Advantages and Improvements:
- Compared to existing XAI-based methods, this tool significantly enhances physicians’ trust in AI recommendations while reducing the cognitive burden of understanding AI’s internal mechanisms.
- Literature evidence provides a perspective familiar and trusted by physicians, bridging the "language gap" between physicians and AI recommendations.
-
Experimental or Evaluation Results:
- A two-phase study involving physicians from various fields confirmed the effectiveness of the design strategy. Physicians acknowledged the practical benefits of combining literature evidence with AI recommendations to improve decision quality.
- Prototype implementation revealed new opportunities: integrating biomedical imaging and genomics evidence; using literature evidence as a standalone decision support tool in resource-limited hospitals.
-
Limitations and Future Directions:
- Current language models may still exhibit errors in processing medical texts. Further research is needed to clarify the tool’s accuracy and summarization capabilities in emergency scenarios.
- Additional validation of this design’s applicability and generalization across other disease domains is required.
- Exploring the integration of more data types (e.g., medical imaging and -omics data) while addressing concerns about potential skill degradation among physicians.
Conclusion
This paper introduces a trust calibration interaction design based on biomedical literature, offering a new direction for improving the explainability of clinical AI and building physician trust. Future research could advance this design concept by enhancing literature integration technologies and exploring cross-domain applicability for broader practical implementation.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can biomedical literature be used to adjust clinicians' trust in AI decision support systems?Category: Medical AI Trust, Clinical Decision Support, and Patient-Provider CollaborationSimilar questionsarrow_forward
- Can biomedical literature more effectively validate applicability and trustworthiness of AI recommendations than traditional AI mechanism explanations?Category: Medical AI Trust, Clinical Decision Support, and Patient-Provider CollaborationSimilar questionsarrow_forward
- What are clinicians' specific needs for evaluating the applicability and robustness of literature support in AI recommendations?Category: Medical AI Trust, Clinical Decision Support, and Patient-Provider CollaborationSimilar questionsarrow_forward
Practical Problems
1- Clinicians struggle to fully trust AI recommendation accuracy, potentially affecting decisions and patient outcomes.Category: Medical AI Trust, Clinical Decision Support, and Patient-Provider CollaborationSimilar questionsarrow_forward
- 100%
“If I Had All the Time in the World”: Ophthalmologists' Perceptions of Anchoring Bias Mitigation in Clinical AI Support
CHI '23· Explainable AI (XAI) +1
- 100%
Rethinking the Role of AI with Physicians in Oncology: Revealing Perspectives from Clinical and Research Workflows
CHI '23· Explainable AI (XAI) +1
- 80%
Human-Centered Tools for Coping with Imperfect Algorithms During Medical Decision-Making
CHI '19· Generative AI (Text, Image, Music, Video) +2
- 75%
Ignore, Trust, or Negotiate: Understanding Clinician Acceptance of AI-Based Treatment Recommendations in Health Care
CHI '23· Explainable AI (XAI) +1
- 67%
Augmenting Pathologists with NaviPath: Design and Evaluation of a Human-AI Collaborative Navigation System
CHI '23· Explainable AI (XAI) +2
- 60%
Designing Theory-Driven User-Centric Explainable AI
CHI '19· Explainable AI (XAI) +1
- 60%
Designing AI for Trust and Collaboration in Time-Constrained Medical Decisions: A Sociotechnical Lens
CHI '21· Explainable AI (XAI) +1
- 60%
Assessing the Impact of Automated Suggestions on Decision Making: Domain Experts Mediate Model Errors but Take Less Initiative
CHI '21· Explainable AI (XAI) +1
- 60%
Healthcare AI Treatment Decision Support: Design Principles to Enhance Clinician Adoption and Trust
CHI '23· Explainable AI (XAI) +1
- 60%
Amplifying Human Capabilities in Prostate Cancer Diagnosis: An Empirical Study of Current Practices and AI Potentials in Radiology
CHI '24· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)