“Do I Trust the AI?” Towards Trustworthy AI-Assisted Diagnosis: Understanding User Perception in LLM-Supported Clinical Reasoning
Authors
Yuchen Wu
ShanghaiTech UniversityPaper Title
“Do I Trust the AI?” Towards Trustworthy AI-Assisted Diagnosis: Understanding User Perception in LLM-Supported Clinical Reasoning
Publication Info
- Topic area: Trust and usability of large language models (LLMs) in clinical reasoning and diagnosis.
- Keywords: Trustworthy AI, large language models, clinical reasoning, human-AI collaboration, diagnostic support, physician perception, AI trust calibration, medical benchmarks, human evaluation, healthcare AI.
Background and Problem
- Problem / challenge: Physicians struggle to accurately perceive and trust the diagnostic capabilities of LLMs due to model opacity and variability in performance. Existing evaluations focus on benchmarks that fail to capture real-world clinical reasoning complexity.
- Significance: Miscalibrated trust in LLMs can lead to either over-reliance or under-reliance, undermining their potential to improve diagnostic outcomes in high-stakes medical contexts.
- Motivation and related work: Prior research has explored human-AI collaboration, trust calibration, and LLM evaluation, but has not sufficiently addressed physicians’ subjective perceptions of LLMs in clinical reasoning. This paper seeks to bridge the gap between benchmark-based evaluations and real-world physician perceptions.
Solution
- Proposed approach: A two-step study to evaluate physicians’ perceptions of LLMs’ clinical reasoning capabilities and compare these perceptions with benchmark-based performance.
- Novelty:
- Development of a structured evaluation framework integrating physicians’ subjective perceptions with objective metrics.
- Introduction of the Perceived Capability Score to quantify LLM diagnostic capabilities from physicians’ perspectives.
- Identification of discrepancies between perception-based and benchmark-based evaluations, highlighting overlooked dimensions in benchmarks.
- Procedure and key techniques:
- Step One: Case Analysis Collection: Nine clinical cases were designed, and analyses were performed by physicians and six LLMs. Diagnostic workflows, including inquiry, diagnosis, and treatment planning, were captured.
- Step Two: Evaluation of Case Analysis: Physicians evaluated the analyses across seven dimensions (e.g., diagnostic accuracy, reasoning soundness, treatment appropriateness). Rankings and scores were used to derive the Perceived Capability Score and compared with benchmark performance.
Results
- Concrete findings:
- LLMs’ perceived capabilities correlated positively but non-linearly with benchmark performance, with diminishing returns at higher benchmark scores.
- Physicians valued dimensions like reasoning coherence and clinical acceptability, which benchmarks often underemphasized.
- Top-performing LLMs outperformed most physicians, except senior specialists, in overall rankings.
- Advantage over baselines: The study revealed that benchmarks focusing solely on diagnostic accuracy fail to capture critical dimensions like evidence acquisition, reasoning coherence, and clinical feasibility, which are crucial for trust calibration.
- Experiments / evaluation:
- Participants: 37 physicians evaluated case analyses across nine clinical cases.
- Metrics: Dimension scores, rankings, and Perceived Capability Scores were analyzed using statistical models (e.g., Bradley–Terry ranking regression, Kendall’s W).
- Benchmarks: Comparison with DiagnosisArena (a clinical diagnosis benchmark).
- Limitations and future work:
- Limited number of LLMs and clinical cases.
- Absence of specialized medical LLMs due to access and resource constraints.
- Future work should involve more diverse cases, advanced virtual patients, and longitudinal studies to refine evaluation frameworks.
Summary
This study investigates physicians’ perceptions of LLMs’ clinical reasoning capabilities and their alignment with benchmark-based evaluations. By introducing the Perceived Capability Score, the authors highlight discrepancies between subjective assessments and traditional benchmarks, emphasizing overlooked dimensions like reasoning coherence and clinical acceptability. The findings underscore the need for multi-dimensional trust calibration and interactive, evidence-based reasoning support to foster effective physician-LLM collaboration. These insights provide a foundation for designing trustworthy AI-assisted diagnostic systems that align with clinical realities.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
Which Contributions Deserve Credit? Perceptions of Attribution in Human-AI Co-Creation
CHI '25· Human-LLM Collaboration +2
- 83%
Co-Disclosing the Computer: LLM-Mediated Computing through Reflective Conversation
CHI '26· Human-LLM Collaboration +2
- 83%
What can AI do for me: Evaluating Machine Learning Interpretations in Cooperative Play
IUI '19· Human-LLM Collaboration +2
- 83%
CAIM: Development and Evaluation of a Cognitive AI Memory Framework for Long-Term Interaction with Intelligent Agents
IUI '26· Human-LLM Collaboration +2
- 83%
User Reliance on AI Support for Collaborative Partner Selection
IUI '26· Human-LLM Collaboration +2
- 71%
Effects of LLM-based Search on Decision Making: Speed, Accuracy, and Overreliance
CHI '25· Human-LLM Collaboration +2
- 71%
More Isn't Always Better: Balancing Decision Accuracy and Conformity Pressures in Multi-AI Advice
CHI '26· Human-LLM Collaboration +2
- 71%
Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving Tasks
CHI '26· Human-LLM Collaboration +2
- 71%
The Impact of Response Latency and Task Type on Human-LLM Interaction and Perception
CHI '26· Human-LLM Collaboration +2
- 71%
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
CHI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)