High Accuracy and Hidden Disparities: Investigating Foundation Model Performance in Clinical Cognitive Assessment

Explainable AI (XAI)AI-Assisted Decision-Making & AutomationMental Health Apps & Online Support CommunitiesChronic Disease Self-Management (Diabetes, Hypertension, etc.)Physicians, Nurses & CliniciansPsychiatrists & PsychotherapistsAI/ML Researchers & Engineers

Paper Title

High Accuracy and Hidden Disparities: Investigating Foundation Model Performance in Clinical Cognitive Assessment

Publication Info

  • Topic area: Evaluation of foundation models in clinical cognitive assessment.
  • Keywords: Foundation models, cognitive assessment, clock drawing test, clinical AI, decomposition framework, human-AI collaboration, healthcare evaluation, anticipatory thinking, visuospatial skills, motor control.

Background and Problem

  • Problem / challenge: Current evaluation metrics for foundation models fail to capture the underlying mechanisms of their decision-making, masking disparities in component-level performance despite high aggregate accuracy.
  • Significance: Accurate cognitive assessments are critical for early detection of conditions like Mild Cognitive Impairment (MCI), which affects 10–20% of people over 65, enabling timely interventions and preserving functional independence.
  • Motivation and related work: Prior research has focused on aggregate metrics like accuracy and sensitivity, which were developed for human clinicians. These metrics fail to account for the differences in how foundation models process information. The clock drawing test (CDT) is a validated tool for cognitive assessment, engaging multiple cognitive domains, making it ideal for evaluating foundation models.

Solution

  • Proposed approach: A theory-driven decomposition framework that breaks down the CDT into 24 binary questions across five cognitive domains to systematically evaluate foundation models.
  • Novelty:
    1. Demonstration that evaluation metrics for human clinicians cannot be directly applied to foundation models.
    2. Development of a decomposition framework to assess specific cognitive operations and their transferability to AI systems.
    3. Discussion of design implications for human-AI collaboration in healthcare, emphasizing evidence-forward systems and abstention as a safety signal.
  • Procedure and key techniques:
    • Decomposed CDT scoring into five cognitive domains: visuospatial skills, motor control, executive function and planning, sustained and selective attention, and anticipatory thinking.
    • Developed 24 binary questions informed by clinical theory and validated through meta-reviews of CDT scoring systems.
    • Conducted a human-AI comparison study using 100 CDT images from the NHATS dataset, with responses collected from three foundation models (ChatGPT, Claude, and Gemini) and human raters.

Results

  • Concrete findings:
    • Models achieved 94% accuracy in dementia screening but only 55% accuracy in granular scoring (0–5 scale).
    • 22.1% error rate in cases where all three models strongly agreed but disagreed with human raters.
    • Performance varied by domain: 88% alignment with humans on rule-based executive function tasks but only 46% on context-based anticipatory thinking tasks.
    • Models abstained three times more than humans (12.4% vs. 4.4%), especially for low-quality images.
  • Advantage over baselines: The decomposition framework revealed hidden disparities in model performance that aggregate metrics obscured, enabling domain-specific insights into model capabilities and limitations.
  • Experiments / evaluation:
    • Used 100 CDT images scored on a validated 6-point scale from the NHATS dataset.
    • Conducted a zero-shot evaluation with structured prompts for foundation models and collected human ratings via Prolific.
    • Analyzed agreement rates, mismatch rates, and abstention patterns across cognitive domains and clock quality levels.
  • Limitations and future work:
    • Small sample size (100 images) and reliance on US-centric NHATS dataset.
    • Prompt sensitivity and dataset bias may affect generalizability.
    • Future work should validate findings across diverse populations and larger datasets.

Summary

This study investigated the performance of foundation models on the clock drawing test (CDT), revealing that high aggregate accuracy masks significant disparities in component-level tasks. Using a theory-driven decomposition framework, the authors demonstrated that models excel in rule-based tasks but struggle with context-dependent judgments like anticipatory thinking. Models abstained three times more than humans, suggesting abstention could serve as a valuable safety signal. The findings highlight the need for AI-specific evaluation metrics and evidence-forward designs in clinical AI systems, emphasizing human oversight and modular outputs for better human-AI collaboration.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222644/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791866
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Explainable AI (XAI), AI-Assisted Decision-Making & Automation, Mental Health Apps & Online Support Communities, Chronic Disease Self-Management (Diabetes, Hypertension, etc.)
work
Professions
Physicians, Nurses & Clinicians, Psychiatrists & Psychotherapists, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
8 related papers