Augmenting Clinical Decision-Making with an Interactive and Interpretable AI Copilot: A Real-World User Study with Clinicians in Nephrology and Obstetrics
Authors
Paper Title
Augmenting Clinical Decision-Making with an Interactive and Interpretable AI Copilot: A Real-World User Study with Clinicians in Nephrology and Obstetrics
Publication Info
- Topic area: Human-AI collaboration in clinical decision-making
- Keywords: Clinical decision support, interpretable AI, human-AI interaction, nephrology, obstetrics, cognitive workload, trust in AI, longitudinal EHR data, LLM-based recommendations, interactive visualization
Background and Problem
- Problem / challenge: Despite advances in AI accuracy for medical tasks, adoption in clinical practice remains limited due to a lack of transparency and alignment with clinicians' reasoning processes. Current systems often act as opaque predictors, bypassing the iterative and evidence-based nature of clinical decision-making.
- Significance: Bridging the "last mile" gap in medical AI could enhance diagnostic workflows, reduce cognitive workload, and improve patient outcomes in high-stakes environments like nephrology and obstetrics.
- Motivation and related work: Previous efforts in Clinical Decision Support Systems (CDSS) have focused on static predictions, often disregarding the need for interpretability and interactive features. Existing studies lack empirical evaluations of fully functional systems in real-world clinical settings and fail to account for differences in expertise among clinicians.
Solution
- Proposed approach: AICare, an interactive and interpretable AI copilot, designed to support clinicians by providing dynamic risk trajectories, feature-level attributions, LLM-driven diagnostic recommendations, and population-level analysis.
- Novelty:
- Introduction of dynamic risk trajectory visualizations to contextualize patient health over time.
- Integration of interactive feature importance tools for granular data interrogation.
- Use of LLMs to generate concise, narrative diagnostic recommendations.
- Empirical evaluation of AICare in real-world settings, highlighting distinct interaction patterns between junior and senior clinicians.
- Procedure and key techniques:
- AICare processes longitudinal EHR data using a bi-GRU and attention mechanisms to handle irregular sampling and dynamic feature attribution.
- Features include interactive visualizations, LLM-generated summaries, and cohort-level comparisons.
- A within-subjects user study with 16 clinicians across nephrology and obstetrics evaluated AICare's impact on efficiency, accuracy, cognitive workload, and trust.
Results
- Concrete findings:
- AICare significantly reduced cognitive workload (p = .023) and increased diagnostic confidence (p = .018).
- Diagnostic accuracy remained comparable to baseline, with no significant overall efficiency gains but distinct behavioral patterns based on expertise.
- Junior clinicians reduced task time by 32% (not statistically significant), while senior clinicians spent more time engaging in adversarial verification.
- Advantage over baselines:
- AICare enabled clinicians to externalize memory-intensive tasks, providing intuitive visualizations and actionable insights.
- Transparency and interactivity fostered trust and supported diverse reasoning styles, unlike static, opaque systems.
- Experiments / evaluation:
- Conducted with 16 clinicians (9 senior, 7 junior) using real-world EHR datasets from nephrology and obstetrics.
- Evaluated using NASA-TLX, SUS, trust scales, and semi-structured interviews.
- Interaction metrics revealed significant differences in verification behaviors between junior and senior clinicians.
- Limitations and future work:
- Modest sample size and reliance on retrospective data limit generalizability.
- Visual complexity and occasional LLM hallucinations identified as areas for improvement.
- Future work includes adaptive interfaces, counterfactual analysis, and longitudinal studies to refine usability and trust calibration.
Summary
AICare demonstrates the potential of interactive and interpretable AI copilots to augment clinical decision-making by reducing cognitive workload and enhancing diagnostic confidence. The system's dynamic visualizations and LLM-driven recommendations support diverse reasoning styles, acting as cognitive scaffolds for novices and verification tools for experts. While transparency fosters trust, it also highlights the need for robust algorithmic reasoning to avoid trust erosion. These findings underscore the importance of designing AI systems that align with clinical workflows and reasoning processes, paving the way for responsible and effective AI integration in healthcare.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 67%
Designing Theory-Driven User-Centric Explainable AI
CHI '19· Explainable AI (XAI) +1
- 67%
Designing AI for Trust and Collaboration in Time-Constrained Medical Decisions: A Sociotechnical Lens
CHI '21· Explainable AI (XAI) +1
- 67%
How Do Users Experience Traceability of AI Systems? Examining Subjective Information Processing Awareness in Automated Insulin Delivery (AID) Systems
IUI '24· Explainable AI (XAI) +1
- 63%
Rapid Assisted Visual Search: Supporting Digital Pathologists with Imperfect AI
IUI '21· EV Charging & Eco-Driving Interfaces +3
Based on Jaccard similarity of research subtopics & professions (≥60%)