Paper Title

Exploring the Future of AI in Clinical Collaboration: A Study on Tumor Board Case Preparation

Publication Info

  • Topic area: Application of AI in high-stakes clinical workflows, specifically multidisciplinary tumor board (MTB) case preparation.
  • Keywords: AI in healthcare, multidisciplinary tumor boards, clinical decision support, large language models, multi-agent systems, trust in AI, clinical workflows, oncology, AI errors, responsible AI.

Background and Problem

  • Problem / challenge: Preparing cases for multidisciplinary tumor boards (MTBs) is time-intensive and complex, requiring clinicians to extract and synthesize information from extensive and unstructured medical records. Existing off-the-shelf AI systems lack the specialization and accuracy needed for high-stakes clinical tasks.
  • Significance: Efficient preparation for MTBs is critical for timely and accurate cancer treatment decisions, directly impacting patient outcomes. Addressing inefficiencies in this process can improve the quality of care and reduce clinician workload.
  • Motivation and related work: Prior research has explored AI for general medical tasks like summarization and documentation, but its application in high-stakes, multidisciplinary workflows remains underexplored. Multi-agent AI frameworks have shown promise in emulating domain-specific reasoning but require further evaluation in real-world clinical contexts.

Solution

  • Proposed approach: The study evaluates two AI systems for MTB preparation: an off-the-shelf assistant (Copilot) and a task-specific multi-agent system (Healthcare Agent Orchestrator, HAO), focusing on their usability, accuracy, and alignment with clinicians’ workflows.
  • Novelty:
    1. Identification of oncologists’ information needs and expectations for AI in MTB preparation.
    2. Analysis of critical AI errors and their propagation into clinical workflows.
    3. Insights into clinicians’ misaligned mental models of AI capabilities.
    4. Design recommendations for task-specific AI systems to better support clinical tasks.
  • Procedure and key techniques:
    1. Conducted a mixed-methods study with 16 oncologists using Copilot and HAO in a randomized A/B testing setup.
    2. Analyzed 163 AI prompts across three task categories: information retrieval, suggestions/options/evidence, and task completion.
    3. Evaluated AI and human-generated case summaries using the TBFact framework.
    4. Collected survey responses and conducted thematic analysis of interviews to assess perceptions and interactions with the AI systems.

Results

  • Concrete findings:
    • HAO achieved higher ratings than Copilot in task confidence (4.62 vs. 3.68) and willingness to use (4.50 vs. 3.56).
    • Four critical errors were identified in Copilot’s responses, including status mix-ups and misinterpretation of clinical decisions. HAO demonstrated no critical errors.
    • Oncologists made four errors in their case summaries, all linked to incorrect AI outputs from Copilot.
  • Advantage over baselines:
    • HAO’s multi-agent structure aligned better with clinicians’ reasoning, providing more personalized and contextually relevant responses.
    • Copilot’s simplicity was appreciated by some but led to critical errors and less nuanced outputs.
  • Experiments / evaluation:
    • Simulated-use study with 16 oncologists preparing two patient cases using both AI systems.
    • Accuracy of AI responses and oncologists’ notes evaluated using TBFact and manual verification.
    • Surveys and thematic analysis assessed usability, trust, and perceptions of AI systems.
  • Limitations and future work:
    • Study focused on pancreatic cancer cases and new patient scenarios, limiting generalizability.
    • Limited female participation (2 out of 16 oncologists).
    • Need for more controlled experiments to isolate the effects of multi-agent architecture.
    • Future work should explore broader use cases, trust calibration techniques, and long-term adoption.

Summary

This study evaluated the use of two AI systems, Copilot and HAO, for preparing patient cases in multidisciplinary tumor board (MTB) meetings. HAO’s task-specific, multi-agent design better aligned with clinicians’ reasoning, achieving higher ratings in confidence and willingness to use, while Copilot exhibited critical errors that propagated into clinical workflows. Oncologists’ misaligned mental models of AI capabilities and the inefficacy of traditional trust-calibration techniques were identified as key challenges. The findings highlight the potential of task-specific AI systems in high-stakes clinical tasks and provide actionable design recommendations to improve their safety, usability, and alignment with clinician workflows.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223465/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790448
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
Honorable Mention
group
Authors
30 authors
sell
Subtopics
Human-LLM Collaboration, Explainable AI (XAI), AI-Assisted Decision-Making & Automation, Mental Health Apps & Online Support Communities
work
Professions
Physicians, Nurses & Clinicians, Psychiatrists & Psychotherapists
article
Content Status
Full text indexed
hub
Related Papers
10 related papers