Open-ended Structured Question Assessment with Human-LLM Collaboration

Human-LLM CollaborationIntelligent Tutoring Systems & Learning AnalyticsParticipatory DesignPrototyping & User TestingUniversity Professors & ResearchersOnline Course Designers

Paper Title

Open-ended Structured Question Assessment with Human-LLM Collaboration

Publication Info

  • Topic area: AI-assisted grading systems for educational assessment
  • Keywords: Open-ended structured questions, large language models, human-AI collaboration, educational assessment, grading efficiency, scoring point analysis, in-context learning, interpretability, iterative refinement, instructor feedback

Background and Problem

  • Problem / challenge: Grading open-ended structured questions (OSQs) is labor-intensive, requiring fine-grained, scoring point–level analysis. Existing automated or human-AI collaborative systems lack support for point-level inspection, correction, and feedback integration.
  • Significance: Accurate and scalable OSQ grading is essential for assessing students’ reasoning and expression, ensuring fairness, and reducing instructors’ workload.
  • Motivation and related work: Prior systems focus on whole-response grading and lack transparency, interpretability, and adaptability to instructor feedback. This paper addresses the gap by introducing a system tailored for fine-grained OSQ grading.

Solution

  • Proposed approach: VeriGrader, a human-AI collaborative system for OSQ grading, integrates LLM-based grading with instructor feedback through an interactive interface.
  • Novelty:
    1. Chain-of-thought prompting and in-context learning for interpretable, point-level grading.
    2. A multi-view interface for efficient verification of response segments, scoring points, and rationales.
    3. Iterative refinement of LLM grading behavior using instructor feedback.
  • Procedure and key techniques:
    • Segment student responses into pieces and map them to predefined scoring points.
    • Classify response pieces as correct, wrong, or unclear, with associated grading reasons.
    • Enable instructors to review, relabel, and confirm grading results through an interactive interface.
    • Use confirmed results as exemplars to iteratively improve LLM performance.

Results

  • Concrete findings:
    • VeriGrader reduced grading time by 56.3% (from 3.82 to 1.67 minutes per OSQ) and improved accuracy by 5.6% (from 86.26% to 91.86%).
    • LLM achieved high standalone accuracy (89.5% for data structure OSQs, 89.2% for numerical methods OSQs).
    • Few-shot exemplars improved accuracy by 2.7% on average, with final human review elevating accuracy to 95%.
  • Advantage over baselines:
    • Outperformed manual grading in both efficiency and accuracy.
    • Improved inter-rater consistency compared to manual grading and LLM-only grading.
  • Experiments / evaluation:
    • User study with 12 participants grading 30 student responses across two OSQs.
    • Metrics: F1-score (accuracy), time per OSQ (efficiency), piece matching rate (PMR), and inter-rater consistency (Fleiss’ Kappa, ICC, Gwet’s AC1).
  • Limitations and future work:
    • Limited applicability to unstructured or multimodal responses (e.g., diagrams, formulas).
    • Challenges in handling ambiguous or information-dense responses.
    • Need for resolving conflicting feedback in multi-instructor workflows.
    • Future directions include multimodal assessment, broader question types, and tolerance-aware evaluation metrics.

Summary

VeriGrader is a human-AI collaborative system designed to address the challenges of grading open-ended structured questions (OSQs). By leveraging LLMs for fine-grained response segmentation and scoring, and integrating instructor feedback through an interactive interface, the system improves grading efficiency, accuracy, and consistency. A user study demonstrated significant performance gains over manual grading, highlighting the system’s potential for scalable and fair educational assessment. Future work will focus on expanding applicability to multimodal and unstructured responses, as well as enhancing multi-instructor workflows.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222211/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791034
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Human-LLM Collaboration, Intelligent Tutoring Systems & Learning Analytics, Participatory Design, Prototyping & User Testing
work
Professions
University Professors & Researchers, Online Course Designers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers