EvaluAId: Human-AI Collaborative Evaluation of Open-Ended Student Essays
Authors
Paper Title
EvaluAId: Human-AI Collaborative Evaluation of Open-Ended Student Essays
Publication Info
- Topic area: Human-AI collaboration in grading open-ended student essays.
- Keywords: Human-AI collaboration, grading, open-ended essays, rubric-based evaluation, feedback synthesis, adaptive benchmarking, automated writing evaluation, evidence identification, educational technology, reflective judgment.
Background and Problem
- Problem / challenge: Open-ended essay grading is challenging due to diverse student submissions, subjective criteria, and the need for personalized feedback. Automated Writing Evaluation (AWE) systems often lack adaptability, prioritize surface-level qualities, and fail to provide transparent, evidence-based grading, reducing grader agency.
- Significance: Accurate and thoughtful grading is critical for student learning and fairness. Scaling grading processes without compromising quality is essential in large classrooms.
- Motivation and related work: Prior work in AWE systems and human-AI collaboration has focused on automation, but these approaches often fail in open-ended, instructor-defined settings. This paper addresses the gap by embedding AI within graders’ reasoning processes to enhance human agency and accountability.
Solution
- Proposed approach: EvaluAId, a human-AI collaborative system that augments graders in rubric-based evaluation of open-ended essays by providing targeted, evidence-first support.
- Novelty:
- Interactive rubric-evidence mapping to help graders locate and justify rubric-aligned evidence.
- Adaptive benchmarking for self-calibration and consistent grading through pairwise comparisons.
- Scaffolded feedback synthesis for personalized, evidence-grounded feedback.
- Human-in-the-loop design to preserve grader control and accountability.
- Procedure and key techniques:
- Decomposes rubric criteria into guiding questions and highlights relevant text in essays.
- Allows graders to set baseline essays for adaptive benchmarking and visualize similarities across submissions.
- Provides AI-generated, editable feedback drafts tied to rubric descriptors and student content.
- Ensures transparency and human oversight by requiring graders to confirm all AI contributions.
Results
- Concrete findings:
- EvaluAId reduced RMSE to 0.744 compared to 1.046 for a conversational AI baseline and 1.031 for an AWE system.
- Achieved higher Pearson correlation with expert scores (0.754 vs. 0.509 for the baseline).
- Produced more actionable feedback with higher ratings for clear directions for improvement (4.042 vs. 3.292 for the baseline).
- Generated feedback drafts with 91% reasoned justifications and 98% actionable suggestions, with a low hallucination rate (7% for justifications, 6% for suggestions).
- Advantage over baselines:
- Improved grading accuracy and inter-rater consistency.
- Higher grader satisfaction, confidence, and perceived collaboration compared to both a conversational AI baseline and an AWE system.
- Feedback was more personalized and aligned with rubric criteria.
- Experiments / evaluation:
- Within-subjects study with 12 TAs comparing EvaluAId to a conversational AI baseline and an AWE system.
- Feedback quality assessed by independent raters across four dimensions (accuracy, criteria-based, improvement directions, tone).
- Semi-structured interviews with 12 TAs, 6 instructors, and 6 students to gather multi-stakeholder perspectives.
- Limitations and future work:
- Limited to text-based grading; multimodal inputs (e.g., diagrams) are not supported.
- Short-term studies; longitudinal deployments are needed to evaluate sustained impact.
- Small sample size and single-course focus limit generalizability.
- Future work includes scaling to diverse disciplines, improving rubric design tools, and addressing grader biases.
Summary
EvaluAId repositions AI as an on-demand collaborator in grading open-ended essays, enhancing human agency through evidence-first workflows. It improves grading accuracy, consistency, and feedback quality by integrating rubric-evidence mapping, adaptive benchmarking, and scaffolded feedback synthesis. EvaluAId outperforms conversational AI and AWE baselines in alignment with expert scores and grader satisfaction. Multi-stakeholder interviews highlight its potential for thoughtful grading and classroom integration, though challenges remain in scaling and addressing transparency. The system advocates for human-in-the-loop evaluation as an alternative to end-to-end automation.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
Expanding a Temporal Vocabulary Towards Designing for “Undiscoverable Learning”
CHI '26· Intelligent Tutoring Systems & Learning Analytics +2
- 83%
'Show It, Don't Just Say It': The Complementary Effects of Instruction Multimodality for Software Guidance
CHI '26· Intelligent Tutoring Systems & Learning Analytics +2
- 71%
Open-ended Structured Question Assessment with Human-LLM Collaboration
CHI '26· Human-LLM Collaboration +3
- 71%
Smarter Together: Enhancing Human-AI Collaborative Grading With Teacher-Cognition Multi-Agent LLM Framework
IUI '26· Human-LLM Collaboration +2
- 67%
Classroom Simulacra: Building Contextual Student Generative Agents in Online Education for Learning Behavioral Simulation
CHI '25· Human-LLM Collaboration +1
- 67%
Good Fences Make Good Learning: How Self-Directed Language Learners Navigate LLM Delegation Decisions
CHI '26· Human-LLM Collaboration +1
- 67%
AskNow: An LLM-powered Interactive System for Real-Time Question Answering in Large-Scale Classrooms
CHI '26· Human-LLM Collaboration +1
- 67%
AI meets Mathematics Education: Supporting Instructors in Large Mathematics Classes with Context-Aware AI
CHI '26· Human-LLM Collaboration +1
- 67%
Can an AI Partner Empower Learners to Ask Critical Questions?
IUI '25· Human-LLM Collaboration +1
- 63%
Codesigning Ripplet: an LLM-Assisted Assessment Authoring System Grounded in a Conceptual Model of Teachers’ Workflows
CHI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)