Accurate Insights, Trustworthy Interactions: Designing a Collaborative AI-Human Multi-Agent System with Knowledge Graph for Diagnosis Prediction

Brain-Computer Interface (BCI) & NeurofeedbackExplainable AI (XAI)AI-Assisted Decision-Making & AutomationPhysicians, Nurses & CliniciansUniversity Professors & ResearchersAI/ML Researchers & Engineers

Research Background and Issues

  • What problems or challenges did the authors identify? Although question-answering systems in the medical domain (MedQA) can support clinical decision-making, traditional systems face difficulties in handling multi-agent interactions and domain-specific knowledge. This leads to reduced accuracy and credibility in clinical decisions. Additionally, large pre-trained language models (PLMs) often make errors or produce inexplicable results when dealing with complex medical cases (the "hallucination" problem), further diminishing user trust in these systems.

  • Why is this issue important? Accurate decision-making in the medical field is critical for patient safety. With the continuous expansion of medical knowledge and the rapid growth of data, clinicians face the challenge of information overload, making it difficult to quickly and comprehensively make accurate decisions. Therefore, developing systems that can provide reliable diagnostic support is essential for improving healthcare quality and reducing diagnostic errors.

  • Research Motivation and Related Work Building on the success of modern pre-trained models in medical information processing, the authors propose a solution to address current issues such as insufficient domain knowledge, lack of transparency, and low credibility. Existing literature suggests that combining knowledge graphs (KG), retrieval-augmented generation (RAG), and multi-agent systems (MAS) can alleviate these challenges to some extent.


Solution

  • What methods or solutions did the authors propose? The authors designed a medical decision support system that integrates multi-agent systems (MAS), medical knowledge graphs (KG), and retrieval-augmented generation (RAG) technology. This system effectively coordinates intelligent agents to process multi-source heterogeneous information, including medical literature, clinical guidelines, and electronic health records, to generate accurate and trustworthy diagnostic recommendations.

  • What are the innovative aspects of this solution?

    1. Multi-Agent System (MAS): Virtual expert agents with clear roles (including retrieval agents, knowledge agents, and coordination agents) simulate the collaborative mechanisms of multidisciplinary doctors.
    2. Knowledge Graph (KG): Provides domain-specific structured knowledge to enhance the interpretability and accuracy of results.
    3. Retrieval-Augmented Generation (RAG): Combines structured and unstructured data to mitigate the "hallucination" problem and ensure generated answers align with medical domain knowledge.
    4. Offers interpretable diagnostic pathways and final reports, enhancing clinicians' trust in the system.
  • What are the implementation steps and key technologies used?

    1. Symptom Summarization: Natural language processing techniques are used to organize user input into structured symptom lists.
    2. Expert Invocation: Based on the summarized symptoms, virtual expert agents from different domains are activated.
    3. Knowledge Retrieval: RAG models retrieve relevant literature and structured knowledge from medical databases and knowledge graphs.
    4. Result Discussion: Collaboration among agents integrates outputs, using weighted averaging or consensus decision models to generate preliminary diagnostic recommendations.
    5. Consensus Building: Multi-criteria decision analysis (MCDA) is employed to generate the final diagnostic report, along with explanatory notes to improve system transparency.

Research Outcomes

  • What specific outcomes were achieved?

    1. The CARE system developed by the authors demonstrated superior performance in diagnostic prediction, excelling in evaluations across multiple medical datasets (MMLU-Med, PubMedQA*, BioASQ).
    2. The system effectively reduced the "hallucination" problem in generative models while enhancing the credibility and interpretability of diagnostic results.
  • What advantages does it have compared to existing solutions? The CARE system outperformed traditional zero-shot models, chain-of-thought (COT) approaches, and models solely using RAG in terms of accuracy, user satisfaction, and information transparency.

    • On the BioASQ dataset, CARE achieved an accuracy of 0.9013, significantly higher than RAG's 0.8948.
    • The CARE system provided richer contextual information and higher interpretability of diagnostic results, enhancing trust among medical professionals.
  • What were the experimental or evaluation results? Through technical evaluations and user studies:

    1. Technical evaluations showed that the CARE system outperformed other methods in accuracy metrics across three datasets.
    2. User studies revealed high satisfaction among participants regarding the system's information relevance, comprehensibility, and usability; improved transparency and privacy protection increased user trust.
    3. Users also suggested future optimization directions, including handling rare diseases, integrating multimodal data (e.g., imaging data), and simplifying operational processes.
  • Limitations and Future Directions

    1. Insufficient rare disease diagnostic capability: Due to the lack of rare disease data, the model performs inadequately in handling complex cases, requiring improvement through specialized datasets or expert knowledge.
    2. Multimodal data processing: The system has limited capabilities in integrating non-text data such as medical imaging and genetic data. Future work should explore standardization and fusion of different data types.
    3. System scalability: As user demands and data volume increase, the system needs optimization to handle large-scale concurrent requests, especially ensuring stability and rapid response in emergency medical scenarios.

Conclusion

This paper proposes an innovative multi-agent medical decision support system that integrates knowledge graphs and retrieval-augmented generation technology to address the shortcomings of traditional systems in accuracy and transparency. Future work should focus on enhancing the system's capabilities in rare disease diagnosis, multimodal data integration, and user experience to provide more reliable and interpretable support for complex decision-making in the medical domain.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188495/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713526
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Brain-Computer Interface (BCI) & Neurofeedback, Explainable AI (XAI), AI-Assisted Decision-Making & Automation
work
Professions
Physicians, Nurses & Clinicians, University Professors & Researchers, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
5 related papers