OpenCD: Empowering Diagnosis of Children's Mathematical Cognition through Open-ended Multimodal Tasks
Authors
Paper Title
OpenCD: Empowering Diagnosis of Children's Mathematical Cognition through Open-ended Multimodal Tasks
Publication Info
- Topic area: Cognitive diagnosis in early mathematics education using AI-assisted multimodal analysis.
- Keywords: Cognitive diagnosis, multimodal tasks, early mathematics, open-ended assessment, Evidence-Centered Design, Vision-Language Models, teacher-AI collaboration, explainable AI, formative assessment, process-based analysis.
Background and Problem
- Problem / challenge: Existing assessment methods in early mathematics education often rely on closed-ended questions, which fail to capture nuanced cognitive processes. Open-ended tasks provide richer insights but are challenging for teachers to analyze systematically due to their unstructured and diverse nature.
- Significance: Understanding children's mathematical cognition is critical for targeted teaching interventions, as foundational deficits in early education can impact long-term learning outcomes.
- Motivation and related work: Prior research highlights the importance of multimodal responses and embodied cognition in mathematics learning. However, current AI systems focus on correctness in structured tasks, neglecting the nuanced cognitive behaviors in open-ended assessments. This paper addresses the gap by designing a system to automate and explain the analysis of multimodal student responses.
Solution
- Proposed approach: OpenCD, an AI system designed to analyze multimodal open-ended tasks and provide teachers with transparent cognitive diagnoses and actionable insights.
- Novelty:
- Integration of Evidence-Centered Design (ECD) to ground AI diagnostics in pedagogical theory.
- Hybrid architecture combining Vision-Language Models (VLMs) and rule-based expert models for interpreting diverse student responses.
- Transparent teacher-facing interface with bi-directional traceability linking diagnostic conclusions to behavioral evidence.
- Class-level analytics for collective instruction alongside individual profiles for personalized intervention.
- Procedure and key techniques:
- Multimodal data (e.g., drawings, object manipulations, speech) is captured and converted into text-and-image scripts for analysis.
- VLMs interpret response processes and detect behaviors, while rule-based modules map these behaviors to cognitive nodes.
- Evidence propagates through a cognitive graph to synthesize holistic diagnoses.
- Diagnoses are visualized in cognitive graphs and comprehensive reports, enabling teachers to trace conclusions back to specific student actions.
Results
- Concrete findings:
- Diagnostic accuracy: 90.3% of AI-generated diagnoses rated as "completely reasonable" by expert teachers.
- Error analysis: Majority of errors stemmed from conservative biases, such as underestimating mastery or misjudging evidence sufficiency.
- User study: Teachers reported reduced mental demand (p=.046) and effort (p=.0045) when using OpenCD, with increased confidence in their diagnoses (p=.0081).
- Advantage over baselines:
- Automates analysis of unstructured multimodal data, addressing scalability challenges in process-based assessment.
- Provides deeper insights into student cognition compared to manual analysis.
- Experiments / evaluation:
- Expert evaluation: Two experienced teachers reviewed diagnoses for 320 cognitive nodes across 16 students.
- User study: 20 teachers analyzed student responses under manual and OpenCD-assisted conditions, with qualitative and quantitative feedback collected.
- Limitations and future work:
- Limited generalizability to higher-grade topics and less digitized contexts.
- Latency in the diagnostic pipeline (~1 minute per response) precludes real-time adaptive questioning.
- Need for scalable methods to construct cognitive graphs and refine intermediate analysis steps.
Summary
OpenCD addresses the challenge of analyzing children's multimodal open-ended responses in early mathematics education by combining Vision-Language Models with rule-based expert systems grounded in Evidence-Centered Design. The system achieves high diagnostic accuracy (90.3%) and significantly reduces teachers' cognitive burden while enhancing their insights into student thinking. Through transparent interfaces and class-level analytics, OpenCD empowers teachers to make informed pedagogical decisions and supports scalable process-based assessment. Future work aims to expand content coverage, optimize latency, and validate long-term impacts on teaching and learning outcomes.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
"Listen to the Teachers": Research-Based Personas for Translating Classroom Realities into Actionable HCI Design
CHI '26· User Research Methods (Interviews, Surveys, Observation) +3
- 86%
OATutor: An Open-source Adaptive Tutoring System and Curated Content Library for Learning Sciences Research
CHI '23· Programming Education & Computational Thinking +2
- 71%
Exploring the Potential of an Intelligent Tutoring System for Sketching Fundamentals
CHI '20· Programming Education & Computational Thinking +1
- 71%
Toward Automated Feedback on Teacher Discourse to Enhance Teacher Learning
CHI '20· Intelligent Tutoring Systems & Learning Analytics +1
- 71%
Expanding a Temporal Vocabulary Towards Designing for “Undiscoverable Learning”
CHI '26· Intelligent Tutoring Systems & Learning Analytics +2
- 71%
'Show It, Don't Just Say It': The Complementary Effects of Instruction Multimodality for Software Guidance
CHI '26· Intelligent Tutoring Systems & Learning Analytics +2
- 63%
Engaging Students with Instructor Solutions in Online Programming Homework
CHI '20· Human-LLM Collaboration +2
- 63%
VizProg: Identifying Misunderstandings by Visualizing Students' Coding Progress
CHI '23· Programming Education & Computational Thinking +1
- 63%
Examining the Role of Peer Acknowledgements on Social Annotations: Unraveling the Psychological Underpinnings
CHI '24· Collaborative Learning & Peer Teaching +2
- 63%
EvaluAId: Human-AI Collaborative Evaluation of Open-Ended Student Essays
CHI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)