High Accuracy and Hidden Disparities: Investigating Foundation Model Performance in Clinical Cognitive Assessment
Authors
Paper Title
High Accuracy and Hidden Disparities: Investigating Foundation Model Performance in Clinical Cognitive Assessment
Publication Info
- Topic area: Evaluation of foundation models in clinical cognitive assessment.
- Keywords: Foundation models, cognitive assessment, clock drawing test, clinical AI, decomposition framework, human-AI collaboration, healthcare evaluation, anticipatory thinking, visuospatial skills, motor control.
Background and Problem
- Problem / challenge: Current evaluation metrics for foundation models fail to capture the underlying mechanisms of their decision-making, masking disparities in component-level performance despite high aggregate accuracy.
- Significance: Accurate cognitive assessments are critical for early detection of conditions like Mild Cognitive Impairment (MCI), which affects 10–20% of people over 65, enabling timely interventions and preserving functional independence.
- Motivation and related work: Prior research has focused on aggregate metrics like accuracy and sensitivity, which were developed for human clinicians. These metrics fail to account for the differences in how foundation models process information. The clock drawing test (CDT) is a validated tool for cognitive assessment, engaging multiple cognitive domains, making it ideal for evaluating foundation models.
Solution
- Proposed approach: A theory-driven decomposition framework that breaks down the CDT into 24 binary questions across five cognitive domains to systematically evaluate foundation models.
- Novelty:
- Demonstration that evaluation metrics for human clinicians cannot be directly applied to foundation models.
- Development of a decomposition framework to assess specific cognitive operations and their transferability to AI systems.
- Discussion of design implications for human-AI collaboration in healthcare, emphasizing evidence-forward systems and abstention as a safety signal.
- Procedure and key techniques:
- Decomposed CDT scoring into five cognitive domains: visuospatial skills, motor control, executive function and planning, sustained and selective attention, and anticipatory thinking.
- Developed 24 binary questions informed by clinical theory and validated through meta-reviews of CDT scoring systems.
- Conducted a human-AI comparison study using 100 CDT images from the NHATS dataset, with responses collected from three foundation models (ChatGPT, Claude, and Gemini) and human raters.
Results
- Concrete findings:
- Models achieved 94% accuracy in dementia screening but only 55% accuracy in granular scoring (0–5 scale).
- 22.1% error rate in cases where all three models strongly agreed but disagreed with human raters.
- Performance varied by domain: 88% alignment with humans on rule-based executive function tasks but only 46% on context-based anticipatory thinking tasks.
- Models abstained three times more than humans (12.4% vs. 4.4%), especially for low-quality images.
- Advantage over baselines: The decomposition framework revealed hidden disparities in model performance that aggregate metrics obscured, enabling domain-specific insights into model capabilities and limitations.
- Experiments / evaluation:
- Used 100 CDT images scored on a validated 6-point scale from the NHATS dataset.
- Conducted a zero-shot evaluation with structured prompts for foundation models and collected human ratings via Prolific.
- Analyzed agreement rates, mismatch rates, and abstention patterns across cognitive domains and clock quality levels.
- Limitations and future work:
- Small sample size (100 images) and reliance on US-centric NHATS dataset.
- Prompt sensitivity and dataset bias may affect generalizability.
- Future work should validate findings across diverse populations and larger datasets.
Summary
This study investigated the performance of foundation models on the clock drawing test (CDT), revealing that high aggregate accuracy masks significant disparities in component-level tasks. Using a theory-driven decomposition framework, the authors demonstrated that models excel in rule-based tasks but struggle with context-dependent judgments like anticipatory thinking. Models abstained three times more than humans, suggesting abstention could serve as a valuable safety signal. The findings highlight the need for AI-specific evaluation metrics and evidence-forward designs in clinical AI systems, emphasizing human oversight and modular outputs for better human-AI collaboration.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
Advancing Patient-Centered Shared Decision-Making with AI Systems for Older Adult Cancer Patients
CHI '24· AI-Assisted Decision-Making & Automation +2
- 63%
Exploring Customizable Interactive Tools for Therapeutic Homework Support in Mental Health Counseling
CHI '26· Human-LLM Collaboration +2
- 63%
Human-centered Perspectives on a Clinical Decision Support System for Intensive Outpatient Veteran PTSD Care
CHI '26· AI-Assisted Decision-Making & Automation +2
- 63%
Digitizing the Pre-consultation Experience: Impacts and Design Recommendations
CHI '26· Human-LLM Collaboration +2
- 63%
Towards Better Health Conversations: The Benefits of Context-seeking
CHI '26· Human-LLM Collaboration +2
- 63%
Exploring the Future of AI in Clinical Collaboration: A Study on Tumor Board Case Preparation
CHI '26· Human-LLM Collaboration +3
- 63%
Accuracy-Time Tradeoffs in AI-Assisted Decision Making under Time Pressure
IUI '24· Explainable AI (XAI) +2
- 63%
Adjust for Trust: Mitigating Trust-Induced Inappropriate Reliance on AI Assistance
IUI '26· Explainable AI (XAI) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)