Smarter Together: Enhancing Human-AI Collaborative Grading With Teacher-Cognition Multi-Agent LLM Framework

Authors

SU

Sanskriti Uma

Geniebook
DD

Dio Dzaky Achmad Mustaqim

Geniebook
Human-LLM CollaborationIntelligent Tutoring Systems & Learning AnalyticsUser Research Methods (Interviews, Surveys, Observation)K-12 TeachersUniversity Professors & ResearchersOnline Course Designers

Automated grading in rubric-based, short-answer open-ended questions often mishandles partial credit, calibration, and actionable transparency, ultimately requiring teachers to reevaluate. This challenge is amplified in resource-constrained settings (e.g., with limited teachers and a large student population), resulting in weaker learning outcomes. To address this challenge, we present the Teacher-Cognition Multi-Agent Grading framework (TC-MAG), which mirrors teachers’ micro-steps via anchored LLM agents for rubric creation, guideline checks, blind double marking, arbitration, and cross-checking to calibrate confidence. Each step produces a concise explanation for targeted review. We first conducted a motivational study to inform the design of TC-MAG. Next, we validated the effectiveness of the TC-MAG framework on a dataset of 2,000 Singapore primary school students’ responses across 1–4-mark mathematics questions with teacher-adjudicated gold labels. TC-MAG attained deployment-level reliability (κ=0.968 on 1-mark; quadratic-weighted κ=0.936 on 2–4 marks) by outperforming human teachers (Δκ=+0.063, p<.001) and state-of-the-art LLM baselines (min Δκ=+0.012, p<.001). In a mixed-methods teacher study (N=14; 12.1 years’ experience), explanation format and TC-MAG’s confidence score influenced whether teachers delegated grading to TC-MAG. Staged explanations yielded greater diagnosticity (LR+ 11.5 vs. 4.60 for summarized explanations), informing a progressive disclosure strategy of explanations based on confidence. Overall, TC-MAG offers replicable multi-agent framework and triage methods for classroom deployment while preserving teacher oversight.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/iui/226570/2026

AdRecommended

Learn AI Coding at CodeNow

At a Glance

Paper Snapshot

fact_check
dataset
Source
IUI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Human-LLM Collaboration, Intelligent Tutoring Systems & Learning Analytics, User Research Methods (Interviews, Surveys, Observation)
work
Professions
K-12 Teachers, University Professors & Researchers, Online Course Designers
article
Content Status
Abstract only
hub
Related Papers
10 related papers