Exploring the Future of AI in Clinical Collaboration: A Study on Tumor Board Case Preparation
Honorable MentionAuthors
Paper Title
Exploring the Future of AI in Clinical Collaboration: A Study on Tumor Board Case Preparation
Publication Info
- Topic area: Application of AI in high-stakes clinical workflows, specifically multidisciplinary tumor board (MTB) case preparation.
- Keywords: AI in healthcare, multidisciplinary tumor boards, clinical decision support, large language models, multi-agent systems, trust in AI, clinical workflows, oncology, AI errors, responsible AI.
Background and Problem
- Problem / challenge: Preparing cases for multidisciplinary tumor boards (MTBs) is time-intensive and complex, requiring clinicians to extract and synthesize information from extensive and unstructured medical records. Existing off-the-shelf AI systems lack the specialization and accuracy needed for high-stakes clinical tasks.
- Significance: Efficient preparation for MTBs is critical for timely and accurate cancer treatment decisions, directly impacting patient outcomes. Addressing inefficiencies in this process can improve the quality of care and reduce clinician workload.
- Motivation and related work: Prior research has explored AI for general medical tasks like summarization and documentation, but its application in high-stakes, multidisciplinary workflows remains underexplored. Multi-agent AI frameworks have shown promise in emulating domain-specific reasoning but require further evaluation in real-world clinical contexts.
Solution
- Proposed approach: The study evaluates two AI systems for MTB preparation: an off-the-shelf assistant (Copilot) and a task-specific multi-agent system (Healthcare Agent Orchestrator, HAO), focusing on their usability, accuracy, and alignment with clinicians’ workflows.
- Novelty:
- Identification of oncologists’ information needs and expectations for AI in MTB preparation.
- Analysis of critical AI errors and their propagation into clinical workflows.
- Insights into clinicians’ misaligned mental models of AI capabilities.
- Design recommendations for task-specific AI systems to better support clinical tasks.
- Procedure and key techniques:
- Conducted a mixed-methods study with 16 oncologists using Copilot and HAO in a randomized A/B testing setup.
- Analyzed 163 AI prompts across three task categories: information retrieval, suggestions/options/evidence, and task completion.
- Evaluated AI and human-generated case summaries using the TBFact framework.
- Collected survey responses and conducted thematic analysis of interviews to assess perceptions and interactions with the AI systems.
Results
- Concrete findings:
- HAO achieved higher ratings than Copilot in task confidence (4.62 vs. 3.68) and willingness to use (4.50 vs. 3.56).
- Four critical errors were identified in Copilot’s responses, including status mix-ups and misinterpretation of clinical decisions. HAO demonstrated no critical errors.
- Oncologists made four errors in their case summaries, all linked to incorrect AI outputs from Copilot.
- Advantage over baselines:
- HAO’s multi-agent structure aligned better with clinicians’ reasoning, providing more personalized and contextually relevant responses.
- Copilot’s simplicity was appreciated by some but led to critical errors and less nuanced outputs.
- Experiments / evaluation:
- Simulated-use study with 16 oncologists preparing two patient cases using both AI systems.
- Accuracy of AI responses and oncologists’ notes evaluated using TBFact and manual verification.
- Surveys and thematic analysis assessed usability, trust, and perceptions of AI systems.
- Limitations and future work:
- Study focused on pancreatic cancer cases and new patient scenarios, limiting generalizability.
- Limited female participation (2 out of 16 oncologists).
- Need for more controlled experiments to isolate the effects of multi-agent architecture.
- Future work should explore broader use cases, trust calibration techniques, and long-term adoption.
Summary
This study evaluated the use of two AI systems, Copilot and HAO, for preparing patient cases in multidisciplinary tumor board (MTB) meetings. HAO’s task-specific, multi-agent design better aligned with clinicians’ reasoning, achieving higher ratings in confidence and willingness to use, while Copilot exhibited critical errors that propagated into clinical workflows. Oncologists’ misaligned mental models of AI capabilities and the inefficacy of traditional trust-calibration techniques were identified as key challenges. The findings highlight the potential of task-specific AI systems in high-stakes clinical tasks and provide actionable design recommendations to improve their safety, usability, and alignment with clinician workflows.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
MindfulDiary: Harnessing Large Language Model to Support Psychiatric Patients' Journaling
CHI '24· Human-LLM Collaboration +2
- 71%
Exploring Customizable Interactive Tools for Therapeutic Homework Support in Mental Health Counseling
CHI '26· Human-LLM Collaboration +2
- 71%
More than Decision Support: Exploring Patients' Longitudinal Usage of Large Language Models in Real-World Healthcare Settings
CHI '26· Human-LLM Collaboration +2
- 71%
Digitizing the Pre-consultation Experience: Impacts and Design Recommendations
CHI '26· Human-LLM Collaboration +2
- 71%
Towards Better Health Conversations: The Benefits of Context-seeking
CHI '26· Human-LLM Collaboration +2
- 67%
Designing Theory-Driven User-Centric Explainable AI
CHI '19· Explainable AI (XAI) +1
- 67%
Designing AI for Trust and Collaboration in Time-Constrained Medical Decisions: A Sociotechnical Lens
CHI '21· Explainable AI (XAI) +1
- 67%
Explainable Notes: Examining How to Unlock Meaning in Medical Notes with Interactivity and Artificial Intelligence
CHI '24· Explainable AI (XAI) +1
- 67%
How Do Users Experience Traceability of AI Systems? Examining Subjective Information Processing Awareness in Automated Insulin Delivery (AID) Systems
IUI '24· Explainable AI (XAI) +1
- 67%
Will Health Experts Adopt a Clinical Decision Support System for Game-Based Digital Biomarkers? Investigating the Impact of Different Explanations on Perceived Ease-of-Use, Perceived Usefulness, and Trust
IUI '25· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)