Document Title

Multimodal Healthcare AI: Identifying and Designing Clinically Relevant Vision-Language Applications for Radiology

Document Information

  • Subject Area: Multimodal healthcare AI, specifically exploring and designing clinically relevant applications of vision-language models (VLMs) in radiology.
  • Keywords: Artificial intelligence, radiology, medical imaging, human-computer interaction, vision-language models, responsible AI, medical workflow.

Research Background and Problem

  • Challenges: Despite significant advancements of multimodal AI models (e.g., vision-language models, VLMs) in natural language processing and image analysis, their clinical applications face challenges such as inconsistent AI performance, lack of trust, and poor workflow integration leading to limited practicality.
  • Importance: Current radiology workflows are complex and demand high efficiency, with pressing needs such as improving radiology report generation, visual search, or query tools to support diagnostic tasks. However, these needs have yet to be fully integrated with AI technologies.
  • Research Motivation: To explore potential applications of VLMs in radiology and investigate how human-centered design approaches can make these technologies better suited for clinical practice, gaining acceptance and adoption by physicians and radiologists.

Solution

  • Research Methodology: The study proposes defining and exploring the potential of vision-language models (VLMs) through human-computer interaction design and participatory design, aligning AI applications with real-world radiology needs.
  • Design Concepts:
    1. VLMs assisting radiologists in generating draft reports (Draft Report Generation).
    2. Enhancing radiologists' ability to review radiology reports (Augmented Report Review).
    3. Providing visual search and text-query tools (Visual Search and Querying).
    4. Developing tools for summarizing patient imaging history (Patient Imaging History Highlights).
  • Key Technologies:
    • Multimodal learning combining images and natural language.
    • Efficiently integrated workflow tools rather than open-ended interactions (e.g., chatbots).
    • Iterative research incorporating feedback from radiologists and physicians to understand workflow adaptability and design requirements.

Research Findings

  • Key Discoveries:

    1. Draft Report Generation: Physician acceptance depends on the precision of AI models; reports must be nearly “perfect” to save time. Radiologists prefer itemized reports rather than long narrative-style outputs.
    2. Augmented Report Review: Physicians value image review functionalities, particularly for locating abnormal areas mentioned in reports. They distrust chat-based Q&A tools and prefer task-specific assistance tools.
    3. Visual Search and Querying: VLMs can enhance diagnostic capabilities by enabling reference comparisons, such as retrieving similar case samples or comparing normal and abnormal imaging features.
    4. Patient Imaging History Highlights: Using VLMs to extract key changes in patient imaging history (e.g., size variations, significant events) can reduce the workload of report screening for physicians.
  • Advantages:

    • Timely assistance in navigating and analyzing complex imaging (e.g., CT scans).
    • Context-specific explanations based on individual patients, reducing repetitive tasks and improving diagnostic efficiency.
    • AI functionalities are confined to “task tools,” minimizing uncertainties associated with open-ended interactions.
  • Experiments and Evaluation:

    • User feedback experiments conducted with 13 radiologists and physicians validated the potential value and feasibility of the four design concepts.
    • Physicians showed greater acceptance of AI tools designed for specific tasks (e.g., image annotation, history summarization) rather than general-purpose intelligent assistants.
  • Limitations and Future Directions:

    1. AI is currently perceived more as an assistive tool, but achieving high performance in more complex tasks remains a challenge.
    2. Additional experiments are needed to test these systems in real-world workflow environments.
    3. Further research is required to develop solutions for user groups such as physician-radiologist interactions and patient report access.

Framework and Output Format

  • The study employs a three-stage methodology: brainstorming the problem space → designing conceptual prototypes → conducting user feedback experiments, providing guidance for designing AI-based medical technologies.
  • Against the backdrop of rapid advancements in multimodal AI technologies, the research focuses on radiology-centered VLM application scenarios, expanding the practical application potential of AI in medical workflows.

Overall Recommendation: The design of multimodal AI systems should prioritize simplifying physicians’ daily workflows rather than generating overly complex or high-risk diagnostic inferences.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/146624/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642013
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
21 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Explainable AI (XAI), Medical & Scientific Data Visualization
work
Professions
Physicians, Nurses & Clinicians, Radiologists & Pathologists
article
Content Status
Full text indexed
hub
Related Papers
5 related papers