“Do I Trust the AI?” Towards Trustworthy AI-Assisted Diagnosis: Understanding User Perception in LLM-Supported Clinical Reasoning

Human-LLM CollaborationExplainable AI (XAI)AI-Assisted Decision-Making & AutomationPhysicians, Nurses & CliniciansAI/ML Researchers & EngineersHCI Researchers

Paper Title

“Do I Trust the AI?” Towards Trustworthy AI-Assisted Diagnosis: Understanding User Perception in LLM-Supported Clinical Reasoning

Publication Info

  • Topic area: Trust and usability of large language models (LLMs) in clinical reasoning and diagnosis.
  • Keywords: Trustworthy AI, large language models, clinical reasoning, human-AI collaboration, diagnostic support, physician perception, AI trust calibration, medical benchmarks, human evaluation, healthcare AI.

Background and Problem

  • Problem / challenge: Physicians struggle to accurately perceive and trust the diagnostic capabilities of LLMs due to model opacity and variability in performance. Existing evaluations focus on benchmarks that fail to capture real-world clinical reasoning complexity.
  • Significance: Miscalibrated trust in LLMs can lead to either over-reliance or under-reliance, undermining their potential to improve diagnostic outcomes in high-stakes medical contexts.
  • Motivation and related work: Prior research has explored human-AI collaboration, trust calibration, and LLM evaluation, but has not sufficiently addressed physicians’ subjective perceptions of LLMs in clinical reasoning. This paper seeks to bridge the gap between benchmark-based evaluations and real-world physician perceptions.

Solution

  • Proposed approach: A two-step study to evaluate physicians’ perceptions of LLMs’ clinical reasoning capabilities and compare these perceptions with benchmark-based performance.
  • Novelty:
    1. Development of a structured evaluation framework integrating physicians’ subjective perceptions with objective metrics.
    2. Introduction of the Perceived Capability Score to quantify LLM diagnostic capabilities from physicians’ perspectives.
    3. Identification of discrepancies between perception-based and benchmark-based evaluations, highlighting overlooked dimensions in benchmarks.
  • Procedure and key techniques:
    1. Step One: Case Analysis Collection: Nine clinical cases were designed, and analyses were performed by physicians and six LLMs. Diagnostic workflows, including inquiry, diagnosis, and treatment planning, were captured.
    2. Step Two: Evaluation of Case Analysis: Physicians evaluated the analyses across seven dimensions (e.g., diagnostic accuracy, reasoning soundness, treatment appropriateness). Rankings and scores were used to derive the Perceived Capability Score and compared with benchmark performance.

Results

  • Concrete findings:
    • LLMs’ perceived capabilities correlated positively but non-linearly with benchmark performance, with diminishing returns at higher benchmark scores.
    • Physicians valued dimensions like reasoning coherence and clinical acceptability, which benchmarks often underemphasized.
    • Top-performing LLMs outperformed most physicians, except senior specialists, in overall rankings.
  • Advantage over baselines: The study revealed that benchmarks focusing solely on diagnostic accuracy fail to capture critical dimensions like evidence acquisition, reasoning coherence, and clinical feasibility, which are crucial for trust calibration.
  • Experiments / evaluation:
    • Participants: 37 physicians evaluated case analyses across nine clinical cases.
    • Metrics: Dimension scores, rankings, and Perceived Capability Scores were analyzed using statistical models (e.g., Bradley–Terry ranking regression, Kendall’s W).
    • Benchmarks: Comparison with DiagnosisArena (a clinical diagnosis benchmark).
  • Limitations and future work:
    • Limited number of LLMs and clinical cases.
    • Absence of specialized medical LLMs due to access and resource constraints.
    • Future work should involve more diverse cases, advanced virtual patients, and longitudinal studies to refine evaluation frameworks.

Summary

This study investigates physicians’ perceptions of LLMs’ clinical reasoning capabilities and their alignment with benchmark-based evaluations. By introducing the Perceived Capability Score, the authors highlight discrepancies between subjective assessments and traditional benchmarks, emphasizing overlooked dimensions like reasoning coherence and clinical acceptability. The findings underscore the need for multi-dimensional trust calibration and interactive, evidence-based reasoning support to foster effective physician-LLM collaboration. These insights provide a foundation for designing trustworthy AI-assisted diagnostic systems that align with clinical realities.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223374/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790835
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
10 authors
sell
Subtopics
Human-LLM Collaboration, Explainable AI (XAI), AI-Assisted Decision-Making & Automation
work
Professions
Physicians, Nurses & Clinicians, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers