Design and Multi-level Evaluation of MAP-X: a Medically Aligned, Patient-Centered AI Explanation System

Explainable AI (XAI)Telemedicine & Remote Patient MonitoringPhysicians, Nurses & CliniciansPsychiatrists & Psychotherapists

Paper Title

Design and Multi-level Evaluation of MAP-X: a Medically Aligned, Patient-Centered AI Explanation System

Publication Info

  • Topic area: Patient-centered AI explanations in healthcare.
  • Keywords: Explainable AI, healthcare AI, patient-centered design, large language models, retrieval-augmented generation, post-stroke speech assessment, clinical validation, multi-level evaluation, trust calibration.

Background and Problem

  • Problem / challenge: Many healthcare AI systems lack tailored explanations for patients, leading to limited transparency, trust, and adoption. Current evaluation frameworks fail to holistically validate explanations across functional, clinician, and patient levels.
  • Significance: Transparent, patient-centered explanations are critical in high-stakes clinical settings to enhance understanding, trust, and shared decision-making, particularly for complex conditions like post-stroke speech disorders.
  • Motivation and related work: Prior research highlights the importance of grounding explanations in clinical reasoning and aligning them with patient needs. However, existing systems often fail to balance clinical validity and patient comprehensibility. Multi-level evaluation frameworks are rare, and most studies focus narrowly on accuracy or anecdotal evidence.

Solution

  • Proposed approach: MAP-X (Medically Aligned, Patient-centered eXplanation), a system that uses retrieval-augmented generation (RAG) and large language models (LLMs) to generate layered, patient-facing explanations grounded in clinical evidence.
  • Novelty:
    1. A Structure–Signal–Source blueprint for designing and validating explanations across clinician-defined structure, model reasoning signal, and clinical data source.
    2. Integration of RAG to reduce hallucinations and ensure explanations are grounded in clinician-authored evidence.
    3. Progressive disclosure interface for layered explanations tailored to patient comprehension.
    4. Multi-phase evaluation framework addressing functional, clinician, and patient perspectives.
  • Procedure and key techniques:
    • Prediction Module: Extracts acoustic features from speech tasks and predicts severity using ensemble models with SHAP values for interpretability.
    • Retrieval Module: Matches patient data to clinically validated reference cases using vector search.
    • Generation Module: Synthesizes retrieved evidence and model outputs into structured explanations using GPT-4.
    • Interface Design: Presents explanations through dashboards, feature-specific insights, and conversational summaries, employing progressive disclosure to manage cognitive load.

Results

  • Concrete findings:
    • Functional evaluation: Macro F1-score of 84.98% for severity classification, subsystem rating MAE of 0.81, and text quality scores of 75.44 (internal relevance) and 71.78 (external accuracy).
    • Clinician evaluation: High ratings for fluency (4.62/5), relevance (4.49/5), and coherence (4.74/5), with partial agreement (2.89/5) in narrative style consistency.
    • Patient evaluation: Trust scores significantly higher for MAP-X (6.39/7) compared to non-explainable control (5.44/7), with a positive trend in explanation satisfaction.
  • Advantage over baselines: MAP-X explanations improved trust and understanding compared to non-explainable controls, while clinicians valued its structured, evidence-grounded outputs as collaborative tools.
  • Experiments / evaluation:
    • Phase 1: Functional evaluation of faithfulness using clinician-defined metrics and rubric-based text scoring.
    • Phase 2: Application-grounded evaluation with 10 speech-language pathologists assessing clinical relevance and workflow fit.
    • Phase 3: End-user evaluation with 15 post-stroke patients comparing MAP-X explanations to non-explainable controls.
  • Limitations and future work:
    • Risk of patient over-trust; need for harm-mitigation protocols.
    • Limited dataset size and reliance on audio-based perceptual ratings.
    • Generalizability restricted to South Korea and chronic stroke survivors.
    • Lack of empirical testing of clinician-mediated workflows and longitudinal impact.

Summary

MAP-X is an AI system designed to generate medically aligned, patient-centered explanations for post-stroke speech assessments. Using a Structure–Signal–Source blueprint and RAG, it produces layered, transparent explanations that align with clinical reasoning and patient needs. A three-phase evaluation demonstrated functional faithfulness, clinician acceptance as a collaborative tool, and improved patient trust and understanding. Limitations include risks of over-trust, dataset constraints, and restricted generalizability. Future work should focus on adaptive explanations, harm mitigation, and longitudinal studies to assess real-world impact.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222180/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790971
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Explainable AI (XAI), Telemedicine & Remote Patient Monitoring
work
Professions
Physicians, Nurses & Clinicians, Psychiatrists & Psychotherapists
article
Content Status
Full text indexed
hub
Related Papers
10 related papers