Prompting, Oversight, and Adoption: Physicians’ Use of Large Language Models for Diagnostic Reasoning in an LMIC

Human-LLM CollaborationAI-Assisted Decision-Making & AutomationAI Ethics, Fairness & AccountabilityTelemedicine & Remote Patient MonitoringPhysicians, Nurses & CliniciansPsychiatrists & PsychotherapistsCommunity Health Workers

Paper Title

Prompting, Oversight, and Adoption: Physicians’ Use of Large Language Models for Diagnostic Reasoning in an LMIC

Publication Info

  • Topic area: Physician interaction with large language models (LLMs) for diagnostic reasoning in low- and middle-income countries (LMICs).
  • Keywords: Large language models, diagnostic reasoning, healthcare AI, LMICs, physician-AI collaboration, prompting strategies, human oversight, cognitive expansion, automation bias, AI adoption.

Background and Problem

  • Problem / challenge: Despite the growing adoption of LLMs in healthcare, there is limited empirical understanding of how clinicians interact with these tools during diagnostic reasoning, especially in LMICs. Key gaps include how physicians prompt, verify, and collaborate with LLMs, and how these interactions influence diagnostic quality and efficiency.
  • Significance: Understanding physician-LLM collaboration is critical for improving diagnostic accuracy, addressing resource constraints, and mitigating risks like automation bias in LMICs, where healthcare systems face acute staffing and infrastructure challenges.
  • Motivation and related work: Previous studies have shown that LLMs can achieve high diagnostic accuracy but often fail to improve physician performance due to suboptimal collaboration. Research in LMICs highlights infrastructural and cultural barriers to AI adoption, but little is known about how clinicians in these settings engage with LLMs in practice. This study aims to address these gaps by focusing on physician interactions with ChatGPT in Pakistan.

Solution

  • Proposed approach: A mixed-methods study combining interaction log analysis and semi-structured interviews to examine how physicians in Pakistan use ChatGPT for diagnostic reasoning.
  • Novelty:
    1. Provides the first detailed analysis of physician-LLM interaction patterns in an LMIC context.
    2. Develops a taxonomy of prompting strategies and oversight mechanisms used by physicians.
    3. Offers design implications for responsible AI integration in resource-constrained healthcare settings.
  • Procedure and key techniques:
    1. Physicians solved six expert-designed clinical vignettes with optional ChatGPT access, logging all interactions.
    2. Diagnostic reasoning scores were evaluated by three licensed physicians.
    3. Semi-structured interviews with 12 participants explored their perceptions, strategies, and challenges regarding AI use.
    4. Data were analyzed to identify prompting strategies, oversight behaviors, and systemic barriers to AI adoption.

Results

  • Concrete findings:
    • Diagnostic accuracy was highest (62.5%) when physicians provided complete clinical context via copy-paste, compared to manual entry (39.1%).
    • Interaction styles varied, with 52.6% of cases showing high reliance on ChatGPT, while others used it for supplementation or not at all.
    • Junior physicians employed structured scaffolding, while seniors used ChatGPT opportunistically for cross-checking.
    • Physicians valued ChatGPT as a "cognitive expander" but expressed concerns about unreliability, privacy, and automation bias.
  • Advantage over baselines: Providing full clinical context significantly improved diagnostic accuracy. Structured prompting and oversight mitigated risks of overreliance and errors.
  • Experiments / evaluation:
    • Nineteen physicians completed six diagnostic vignettes, with diagnostic reasoning scores evaluated by three experts.
    • Interaction logs and interviews revealed diverse prompting strategies, oversight behaviors, and adoption barriers.
    • Quantitative analysis showed significant variation in diagnostic accuracy by case and input method.
  • Limitations and future work:
    • Small sample size (N=19 for logs, N=12 for interviews) limits generalizability.
    • Study focused on AI-trained physicians in urban settings with stable internet, which may not reflect broader LMIC contexts.
    • Future work should explore longitudinal adoption, patient outcomes, and contextual barriers in rural or low-resource settings.

Summary

This study provides an in-depth analysis of how physicians in Pakistan interact with ChatGPT for diagnostic reasoning. It identifies diverse prompting strategies, such as diagnostic structuring, adversarial testing, and efficiency-driven shortcuts, while emphasizing the importance of human oversight to maintain diagnostic accuracy. Physicians valued ChatGPT as a cognitive aid but expressed concerns about reliability, privacy, and systemic barriers to adoption. The findings highlight the need for context-sensitive AI design, including structured prompting templates, seamless integration with electronic health records, and tailored interfaces for different experience levels. These insights inform the responsible deployment of AI in resource-constrained healthcare environments.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223312/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791761
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
8 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation, AI Ethics, Fairness & Accountability, Telemedicine & Remote Patient Monitoring
work
Professions
Physicians, Nurses & Clinicians, Psychiatrists & Psychotherapists, Community Health Workers
article
Content Status
Full text indexed
hub
Related Papers
6 related papers