Prompting, Oversight, and Adoption: Physicians’ Use of Large Language Models for Diagnostic Reasoning in an LMIC
Authors
Muhammad Hamad Alizai
LUMSPaper Title
Prompting, Oversight, and Adoption: Physicians’ Use of Large Language Models for Diagnostic Reasoning in an LMIC
Publication Info
- Topic area: Physician interaction with large language models (LLMs) for diagnostic reasoning in low- and middle-income countries (LMICs).
- Keywords: Large language models, diagnostic reasoning, healthcare AI, LMICs, physician-AI collaboration, prompting strategies, human oversight, cognitive expansion, automation bias, AI adoption.
Background and Problem
- Problem / challenge: Despite the growing adoption of LLMs in healthcare, there is limited empirical understanding of how clinicians interact with these tools during diagnostic reasoning, especially in LMICs. Key gaps include how physicians prompt, verify, and collaborate with LLMs, and how these interactions influence diagnostic quality and efficiency.
- Significance: Understanding physician-LLM collaboration is critical for improving diagnostic accuracy, addressing resource constraints, and mitigating risks like automation bias in LMICs, where healthcare systems face acute staffing and infrastructure challenges.
- Motivation and related work: Previous studies have shown that LLMs can achieve high diagnostic accuracy but often fail to improve physician performance due to suboptimal collaboration. Research in LMICs highlights infrastructural and cultural barriers to AI adoption, but little is known about how clinicians in these settings engage with LLMs in practice. This study aims to address these gaps by focusing on physician interactions with ChatGPT in Pakistan.
Solution
- Proposed approach: A mixed-methods study combining interaction log analysis and semi-structured interviews to examine how physicians in Pakistan use ChatGPT for diagnostic reasoning.
- Novelty:
- Provides the first detailed analysis of physician-LLM interaction patterns in an LMIC context.
- Develops a taxonomy of prompting strategies and oversight mechanisms used by physicians.
- Offers design implications for responsible AI integration in resource-constrained healthcare settings.
- Procedure and key techniques:
- Physicians solved six expert-designed clinical vignettes with optional ChatGPT access, logging all interactions.
- Diagnostic reasoning scores were evaluated by three licensed physicians.
- Semi-structured interviews with 12 participants explored their perceptions, strategies, and challenges regarding AI use.
- Data were analyzed to identify prompting strategies, oversight behaviors, and systemic barriers to AI adoption.
Results
- Concrete findings:
- Diagnostic accuracy was highest (62.5%) when physicians provided complete clinical context via copy-paste, compared to manual entry (39.1%).
- Interaction styles varied, with 52.6% of cases showing high reliance on ChatGPT, while others used it for supplementation or not at all.
- Junior physicians employed structured scaffolding, while seniors used ChatGPT opportunistically for cross-checking.
- Physicians valued ChatGPT as a "cognitive expander" but expressed concerns about unreliability, privacy, and automation bias.
- Advantage over baselines: Providing full clinical context significantly improved diagnostic accuracy. Structured prompting and oversight mitigated risks of overreliance and errors.
- Experiments / evaluation:
- Nineteen physicians completed six diagnostic vignettes, with diagnostic reasoning scores evaluated by three experts.
- Interaction logs and interviews revealed diverse prompting strategies, oversight behaviors, and adoption barriers.
- Quantitative analysis showed significant variation in diagnostic accuracy by case and input method.
- Limitations and future work:
- Small sample size (N=19 for logs, N=12 for interviews) limits generalizability.
- Study focused on AI-trained physicians in urban settings with stable internet, which may not reflect broader LMIC contexts.
- Future work should explore longitudinal adoption, patient outcomes, and contextual barriers in rural or low-resource settings.
Summary
This study provides an in-depth analysis of how physicians in Pakistan interact with ChatGPT for diagnostic reasoning. It identifies diverse prompting strategies, such as diagnostic structuring, adversarial testing, and efficiency-driven shortcuts, while emphasizing the importance of human oversight to maintain diagnostic accuracy. Physicians valued ChatGPT as a cognitive aid but expressed concerns about reliability, privacy, and systemic barriers to adoption. The findings highlight the need for context-sensitive AI design, including structured prompting templates, seamless integration with electronic health records, and tailored interfaces for different experience levels. These insights inform the responsible deployment of AI in resource-constrained healthcare environments.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 63%
Promise or Peril? Exploring Black Adults' Perspectives on the Use of Artificial Intelligence in Health Contexts
CHI '26· AI Ethics, Fairness & Accountability +2
- 63%
Words to Describe What I’m Feeling: Exploring the Potential of AI Agents for High Subjectivity Decisions in Advance Care Planning
CHI '26· AI-Assisted Decision-Making & Automation +2
- 63%
More than Decision Support: Exploring Patients' Longitudinal Usage of Large Language Models in Real-World Healthcare Settings
CHI '26· Human-LLM Collaboration +2
- 63%
Large Language Models in Peer-Run Community Behavioral Health Services: Understanding Peer Specialists and Service Users’ Perspectives on Opportunities, Risks, and Mitigation Strategies
CHI '26· Human-LLM Collaboration +2
- 63%
With, not For: Co-Designing a Patient-Facing AI Companion Concept for the Emergency Department Waiting Area
CHI '26· AI-Assisted Decision-Making & Automation +2
- 63%
Who Does What? Archetypes of Roles Assigned to LLMs During Human-AI Decision-Making
CHI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)