EvaluAId: Human-AI Collaborative Evaluation of Open-Ended Student Essays

Human-LLM CollaborationIntelligent Tutoring Systems & Learning AnalyticsUser Research Methods (Interviews, Surveys, Observation)Prototyping & User TestingUniversity Professors & ResearchersOnline Course Designers

Paper Title

EvaluAId: Human-AI Collaborative Evaluation of Open-Ended Student Essays

Publication Info

  • Topic area: Human-AI collaboration in grading open-ended student essays.
  • Keywords: Human-AI collaboration, grading, open-ended essays, rubric-based evaluation, feedback synthesis, adaptive benchmarking, automated writing evaluation, evidence identification, educational technology, reflective judgment.

Background and Problem

  • Problem / challenge: Open-ended essay grading is challenging due to diverse student submissions, subjective criteria, and the need for personalized feedback. Automated Writing Evaluation (AWE) systems often lack adaptability, prioritize surface-level qualities, and fail to provide transparent, evidence-based grading, reducing grader agency.
  • Significance: Accurate and thoughtful grading is critical for student learning and fairness. Scaling grading processes without compromising quality is essential in large classrooms.
  • Motivation and related work: Prior work in AWE systems and human-AI collaboration has focused on automation, but these approaches often fail in open-ended, instructor-defined settings. This paper addresses the gap by embedding AI within graders’ reasoning processes to enhance human agency and accountability.

Solution

  • Proposed approach: EvaluAId, a human-AI collaborative system that augments graders in rubric-based evaluation of open-ended essays by providing targeted, evidence-first support.
  • Novelty:
    1. Interactive rubric-evidence mapping to help graders locate and justify rubric-aligned evidence.
    2. Adaptive benchmarking for self-calibration and consistent grading through pairwise comparisons.
    3. Scaffolded feedback synthesis for personalized, evidence-grounded feedback.
    4. Human-in-the-loop design to preserve grader control and accountability.
  • Procedure and key techniques:
    • Decomposes rubric criteria into guiding questions and highlights relevant text in essays.
    • Allows graders to set baseline essays for adaptive benchmarking and visualize similarities across submissions.
    • Provides AI-generated, editable feedback drafts tied to rubric descriptors and student content.
    • Ensures transparency and human oversight by requiring graders to confirm all AI contributions.

Results

  • Concrete findings:
    • EvaluAId reduced RMSE to 0.744 compared to 1.046 for a conversational AI baseline and 1.031 for an AWE system.
    • Achieved higher Pearson correlation with expert scores (0.754 vs. 0.509 for the baseline).
    • Produced more actionable feedback with higher ratings for clear directions for improvement (4.042 vs. 3.292 for the baseline).
    • Generated feedback drafts with 91% reasoned justifications and 98% actionable suggestions, with a low hallucination rate (7% for justifications, 6% for suggestions).
  • Advantage over baselines:
    • Improved grading accuracy and inter-rater consistency.
    • Higher grader satisfaction, confidence, and perceived collaboration compared to both a conversational AI baseline and an AWE system.
    • Feedback was more personalized and aligned with rubric criteria.
  • Experiments / evaluation:
    • Within-subjects study with 12 TAs comparing EvaluAId to a conversational AI baseline and an AWE system.
    • Feedback quality assessed by independent raters across four dimensions (accuracy, criteria-based, improvement directions, tone).
    • Semi-structured interviews with 12 TAs, 6 instructors, and 6 students to gather multi-stakeholder perspectives.
  • Limitations and future work:
    • Limited to text-based grading; multimodal inputs (e.g., diagrams) are not supported.
    • Short-term studies; longitudinal deployments are needed to evaluate sustained impact.
    • Small sample size and single-course focus limit generalizability.
    • Future work includes scaling to diverse disciplines, improving rubric design tools, and addressing grader biases.

Summary

EvaluAId repositions AI as an on-demand collaborator in grading open-ended essays, enhancing human agency through evidence-first workflows. It improves grading accuracy, consistency, and feedback quality by integrating rubric-evidence mapping, adaptive benchmarking, and scaffolded feedback synthesis. EvaluAId outperforms conversational AI and AWE baselines in alignment with expert scores and grader satisfaction. Multi-stakeholder interviews highlight its potential for thoughtful grading and classroom integration, though challenges remain in scaling and addressing transparency. The system advocates for human-in-the-loop evaluation as an alternative to end-to-end automation.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/221853/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790814
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Human-LLM Collaboration, Intelligent Tutoring Systems & Learning Analytics, User Research Methods (Interviews, Surveys, Observation), Prototyping & User Testing
work
Professions
University Professors & Researchers, Online Course Designers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers