Smarter Together: Enhancing Human-AI Collaborative Grading With Teacher-Cognition Multi-Agent LLM Framework
Authors
Automated grading in rubric-based, short-answer open-ended questions often mishandles partial credit, calibration, and actionable transparency, ultimately requiring teachers to reevaluate. This challenge is amplified in resource-constrained settings (e.g., with limited teachers and a large student population), resulting in weaker learning outcomes. To address this challenge, we present the Teacher-Cognition Multi-Agent Grading framework (TC-MAG), which mirrors teachers’ micro-steps via anchored LLM agents for rubric creation, guideline checks, blind double marking, arbitration, and cross-checking to calibrate confidence. Each step produces a concise explanation for targeted review. We first conducted a motivational study to inform the design of TC-MAG. Next, we validated the effectiveness of the TC-MAG framework on a dataset of 2,000 Singapore primary school students’ responses across 1–4-mark mathematics questions with teacher-adjudicated gold labels. TC-MAG attained deployment-level reliability (κ=0.968 on 1-mark; quadratic-weighted κ=0.936 on 2–4 marks) by outperforming human teachers (Δκ=+0.063, p<.001) and state-of-the-art LLM baselines (min Δκ=+0.012, p<.001). In a mixed-methods teacher study (N=14; 12.1 years’ experience), explanation format and TC-MAG’s confidence score influenced whether teachers delegated grading to TC-MAG. Staged explanations yielded greater diagnosticity (LR+ 11.5 vs. 4.60 for summarized explanations), informing a progressive disclosure strategy of explanations based on confidence. Overall, TC-MAG offers replicable multi-agent framework and triage methods for classroom deployment while preserving teacher oversight.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
Toward Automated Feedback on Teacher Discourse to Enhance Teacher Learning
CHI '20· Intelligent Tutoring Systems & Learning Analytics +1
- 83%
Charting the Future of AI in Project-Based Learning: A Co-Design Exploration with Students
CHI '24· Human-LLM Collaboration +1
- 83%
Unlocking Scientific Concepts: How Effective Are LLM-Generated Analogies for Student Understanding and Classroom Practice?
CHI '25· Human-LLM Collaboration +1
- 71%
OATutor: An Open-source Adaptive Tutoring System and Curated Content Library for Learning Sciences Research
CHI '23· Programming Education & Computational Thinking +2
- 71%
ClassMeta: Designing Interactive Virtual Classmate to Promote VR Classroom Participation
CHI '24· Social & Collaborative VR +2
- 71%
VIVID: Human-AI Collaborative Authoring of Vicarious Dialogues from Lecture Videos
CHI '24· Human-LLM Collaboration +2
- 71%
TutorCraftEase: Enhancing Pedagogical Question Creation with Large Language Models
CHI '25· Human-LLM Collaboration +2
- 71%
TeachTune: Reviewing Pedagogical Agents Against Diverse Student Profiles with Simulated Students
CHI '25· Generative AI (Text, Image, Music, Video) +2
- 71%
EvaluAId: Human-AI Collaborative Evaluation of Open-Ended Student Essays
CHI '26· Human-LLM Collaboration +3
- 71%
Exploring Teacher-Chatbot Interaction and Affect in Block-Based Programming
CHI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)