Design and Multi-level Evaluation of MAP-X: a Medically Aligned, Patient-Centered AI Explanation System
Authors
Paper Title
Design and Multi-level Evaluation of MAP-X: a Medically Aligned, Patient-Centered AI Explanation System
Publication Info
- Topic area: Patient-centered AI explanations in healthcare.
- Keywords: Explainable AI, healthcare AI, patient-centered design, large language models, retrieval-augmented generation, post-stroke speech assessment, clinical validation, multi-level evaluation, trust calibration.
Background and Problem
- Problem / challenge: Many healthcare AI systems lack tailored explanations for patients, leading to limited transparency, trust, and adoption. Current evaluation frameworks fail to holistically validate explanations across functional, clinician, and patient levels.
- Significance: Transparent, patient-centered explanations are critical in high-stakes clinical settings to enhance understanding, trust, and shared decision-making, particularly for complex conditions like post-stroke speech disorders.
- Motivation and related work: Prior research highlights the importance of grounding explanations in clinical reasoning and aligning them with patient needs. However, existing systems often fail to balance clinical validity and patient comprehensibility. Multi-level evaluation frameworks are rare, and most studies focus narrowly on accuracy or anecdotal evidence.
Solution
- Proposed approach: MAP-X (Medically Aligned, Patient-centered eXplanation), a system that uses retrieval-augmented generation (RAG) and large language models (LLMs) to generate layered, patient-facing explanations grounded in clinical evidence.
- Novelty:
- A Structure–Signal–Source blueprint for designing and validating explanations across clinician-defined structure, model reasoning signal, and clinical data source.
- Integration of RAG to reduce hallucinations and ensure explanations are grounded in clinician-authored evidence.
- Progressive disclosure interface for layered explanations tailored to patient comprehension.
- Multi-phase evaluation framework addressing functional, clinician, and patient perspectives.
- Procedure and key techniques:
- Prediction Module: Extracts acoustic features from speech tasks and predicts severity using ensemble models with SHAP values for interpretability.
- Retrieval Module: Matches patient data to clinically validated reference cases using vector search.
- Generation Module: Synthesizes retrieved evidence and model outputs into structured explanations using GPT-4.
- Interface Design: Presents explanations through dashboards, feature-specific insights, and conversational summaries, employing progressive disclosure to manage cognitive load.
Results
- Concrete findings:
- Functional evaluation: Macro F1-score of 84.98% for severity classification, subsystem rating MAE of 0.81, and text quality scores of 75.44 (internal relevance) and 71.78 (external accuracy).
- Clinician evaluation: High ratings for fluency (4.62/5), relevance (4.49/5), and coherence (4.74/5), with partial agreement (2.89/5) in narrative style consistency.
- Patient evaluation: Trust scores significantly higher for MAP-X (6.39/7) compared to non-explainable control (5.44/7), with a positive trend in explanation satisfaction.
- Advantage over baselines: MAP-X explanations improved trust and understanding compared to non-explainable controls, while clinicians valued its structured, evidence-grounded outputs as collaborative tools.
- Experiments / evaluation:
- Phase 1: Functional evaluation of faithfulness using clinician-defined metrics and rubric-based text scoring.
- Phase 2: Application-grounded evaluation with 10 speech-language pathologists assessing clinical relevance and workflow fit.
- Phase 3: End-user evaluation with 15 post-stroke patients comparing MAP-X explanations to non-explainable controls.
- Limitations and future work:
- Risk of patient over-trust; need for harm-mitigation protocols.
- Limited dataset size and reliance on audio-based perceptual ratings.
- Generalizability restricted to South Korea and chronic stroke survivors.
- Lack of empirical testing of clinician-mediated workflows and longitudinal impact.
Summary
MAP-X is an AI system designed to generate medically aligned, patient-centered explanations for post-stroke speech assessments. Using a Structure–Signal–Source blueprint and RAG, it produces layered, transparent explanations that align with clinical reasoning and patient needs. A three-phase evaluation demonstrated functional faithfulness, clinician acceptance as a collaborative tool, and improved patient trust and understanding. Limitations include risks of over-trust, dataset constraints, and restricted generalizability. Future work should focus on adaptive explanations, harm mitigation, and longitudinal studies to assess real-world impact.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 67%
DiagLink: A Dual-User Diagnostic Assistance System by Synergizing Experts with LLMs and Knowledge Graphs
CHI '26· Human-LLM Collaboration +2
- 67%
Uncertainty and Risk at the Point of Care: Implications of Patient-Generated ECGs and Algorithmic Interpretations for Clinical Decision Making
CHI '26· Biosensors & Physiological Monitoring +2
- 60%
Designing Theory-Driven User-Centric Explainable AI
CHI '19· Explainable AI (XAI) +1
- 60%
OralCam: Enabling Self-Examination and Awareness of Oral Health Using a Smartphone Camera
CHI '20· Mental Health Apps & Online Support Communities +1
- 60%
Designing AI for Trust and Collaboration in Time-Constrained Medical Decisions: A Sociotechnical Lens
CHI '21· Explainable AI (XAI) +1
- 60%
Leveraging Implementation Science in Human-Centred Design for Digital Health
CHI '24· Mental Health Apps & Online Support Communities +1
- 60%
Explainable Notes: Examining How to Unlock Meaning in Medical Notes with Interactivity and Artificial Intelligence
CHI '24· Explainable AI (XAI) +1
- 60%
``It Is a Moving Process'': Understanding the Evolution of Explainability Needs of Clinicians in Pulmonary Medicine
CHI '24· Explainable AI (XAI) +1
- 60%
Human-Centered Personalization in Radiology AI: Evaluating Trust, Usability, and Cross-Hospital Robustness
CHI '26· Explainable AI (XAI) +1
- 60%
How Do Users Experience Traceability of AI Systems? Examining Subjective Information Processing Awareness in Automated Insulin Delivery (AID) Systems
IUI '24· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)