Beyond Scores: Explainable Intelligent Assessment Strengthens Pre-service Teachers' Assessment Literacy

Explainable AI (XAI)Intelligent Tutoring Systems & Learning AnalyticsK-12 TeachersUniversity Professors & ResearchersOnline Course Designers

Paper Title

Beyond Scores: Explainable Intelligent Assessment Strengthens Pre-service Teachers' Assessment Literacy

Publication Info

  • Topic area: Explainable AI in education, focusing on teacher assessment literacy.
  • Keywords: Explainable AI, assessment literacy, cognitive diagnostic assessment, teacher education, reflection, self-regulation, instructional decision-making, personalized learning, diagnostic reasoning, educational technology.

Background and Problem

  • Problem / challenge: Pre-service teachers struggle to interpret and act on complex assessment data due to opaque outputs and lack of scaffolding in existing tools, hindering their development of assessment literacy (AL).
  • Significance: Strengthening AL is critical for effective personalized education, enabling teachers to make evidence-based instructional decisions and support student learning.
  • Motivation and related work: Prior work on cognitive diagnostic assessments (CDAs) and explainable AI (XAI) has focused on algorithmic transparency or learner-facing feedback but has largely neglected teacher-facing tools that align with classroom workflows. Existing CDA tools often present technical outputs as fixed endpoints, limiting their practical utility for teachers.

Solution

  • Proposed approach: XIA (eXplainable Intelligent Assessment), a platform that integrates visualized diagnostic reasoning with contrastive and counterfactual explanations to scaffold pre-service teachers’ assessment literacy.
  • Novelty:
    1. Design requirements and principles for teacher-facing explainable assessment tools, emphasizing clarity, sufficiency, and actionability.
    2. A system architecture combining cognitive diagnostic modeling, explanation generation, and teacher-centric interfaces.
    3. Empirical evidence linking explanatory scaffolding to improvements in reflection, self-regulation, and assessment awareness.
  • Procedure and key techniques:
    • Teachers upload student test data, which is processed into decision-support information (e.g., test summaries, question-level analysis).
    • Diagnostic reasoning and explanation modules provide visualized reasoning paths, contrastive explanations (e.g., "why result P rather than Q"), and counterfactual explanations (e.g., "what if mastery assumptions changed").
    • Two interfaces: (1) Instructional Decision-Support Interface for statistical insights and (2) Diagnostic Reasoning and Explanation Interface for deeper reflection and calibration.

Results

  • Concrete findings:
    • Full-support group (accessing both interfaces) showed significant improvements in reflection (ΔM = 0.86, p < 0.001), self-regulation (ΔM = 0.93, p < 0.01), and assessment awareness (ΔM = 0.68, p < 0.01).
    • Mean Absolute Error (MAE) in assessment accuracy reduced significantly in the full-support group (ΔM = -0.063, p = 0.009).
  • Advantage over baselines:
    • Full-support group outperformed both the control group and decision-support group in assessment awareness and error reduction.
    • Decision-support group showed moderate gains in reflection and self-regulation but lacked systematic improvements in assessment accuracy.
  • Experiments / evaluation:
    • Mixed-methods study with 21 pre-service teachers divided into three groups: control (no tool), decision-support, and full-support.
    • Pre- and post-tests measured assessment accuracy and AL dimensions (reflection, self-regulation, assessment awareness).
    • Semi-structured interviews revealed shifts from score-based to evidence-based reasoning.
  • Limitations and future work:
    • Short-term, single-session design limits insights into long-term impacts.
    • Focused on mathematics/technology education; generalizability to other disciplines or in-service teachers is unclear.
    • Future work should explore longitudinal designs, multimodal data, and disentangle the effects of different explanatory mechanisms.

Summary

This study introduces XIA, an explainable intelligent assessment platform designed to enhance pre-service teachers’ assessment literacy by integrating diagnostic reasoning with contrastive and counterfactual explanations. Results from a controlled study indicate that XIA supports significant improvements in reflection, self-regulation, and assessment awareness, with the full-support version showing the strongest gains. The findings highlight the potential of explainable AI to scaffold evidence-based reasoning and instructional decision-making in teacher education. Future research should extend this approach to diverse contexts and examine its long-term impact in authentic classroom settings.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222097/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791230
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Explainable AI (XAI), Intelligent Tutoring Systems & Learning Analytics
work
Professions
K-12 Teachers, University Professors & Researchers, Online Course Designers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers