Open-ended Structured Question Assessment with Human-LLM Collaboration
Authors
Paper Title
Open-ended Structured Question Assessment with Human-LLM Collaboration
Publication Info
- Topic area: AI-assisted grading systems for educational assessment
- Keywords: Open-ended structured questions, large language models, human-AI collaboration, educational assessment, grading efficiency, scoring point analysis, in-context learning, interpretability, iterative refinement, instructor feedback
Background and Problem
- Problem / challenge: Grading open-ended structured questions (OSQs) is labor-intensive, requiring fine-grained, scoring point–level analysis. Existing automated or human-AI collaborative systems lack support for point-level inspection, correction, and feedback integration.
- Significance: Accurate and scalable OSQ grading is essential for assessing students’ reasoning and expression, ensuring fairness, and reducing instructors’ workload.
- Motivation and related work: Prior systems focus on whole-response grading and lack transparency, interpretability, and adaptability to instructor feedback. This paper addresses the gap by introducing a system tailored for fine-grained OSQ grading.
Solution
- Proposed approach: VeriGrader, a human-AI collaborative system for OSQ grading, integrates LLM-based grading with instructor feedback through an interactive interface.
- Novelty:
- Chain-of-thought prompting and in-context learning for interpretable, point-level grading.
- A multi-view interface for efficient verification of response segments, scoring points, and rationales.
- Iterative refinement of LLM grading behavior using instructor feedback.
- Procedure and key techniques:
- Segment student responses into pieces and map them to predefined scoring points.
- Classify response pieces as correct, wrong, or unclear, with associated grading reasons.
- Enable instructors to review, relabel, and confirm grading results through an interactive interface.
- Use confirmed results as exemplars to iteratively improve LLM performance.
Results
- Concrete findings:
- VeriGrader reduced grading time by 56.3% (from 3.82 to 1.67 minutes per OSQ) and improved accuracy by 5.6% (from 86.26% to 91.86%).
- LLM achieved high standalone accuracy (89.5% for data structure OSQs, 89.2% for numerical methods OSQs).
- Few-shot exemplars improved accuracy by 2.7% on average, with final human review elevating accuracy to 95%.
- Advantage over baselines:
- Outperformed manual grading in both efficiency and accuracy.
- Improved inter-rater consistency compared to manual grading and LLM-only grading.
- Experiments / evaluation:
- User study with 12 participants grading 30 student responses across two OSQs.
- Metrics: F1-score (accuracy), time per OSQ (efficiency), piece matching rate (PMR), and inter-rater consistency (Fleiss’ Kappa, ICC, Gwet’s AC1).
- Limitations and future work:
- Limited applicability to unstructured or multimodal responses (e.g., diagrams, formulas).
- Challenges in handling ambiguous or information-dense responses.
- Need for resolving conflicting feedback in multi-instructor workflows.
- Future directions include multimodal assessment, broader question types, and tolerance-aware evaluation metrics.
Summary
VeriGrader is a human-AI collaborative system designed to address the challenges of grading open-ended structured questions (OSQs). By leveraging LLMs for fine-grained response segmentation and scoring, and integrating instructor feedback through an interactive interface, the system improves grading efficiency, accuracy, and consistency. A user study demonstrated significant performance gains over manual grading, highlighting the system’s potential for scalable and fair educational assessment. Future work will focus on expanding applicability to multimodal and unstructured responses, as well as enhancing multi-instructor workflows.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 86%
Codesigning Ripplet: an LLM-Assisted Assessment Authoring System Grounded in a Conceptual Model of Teachers’ Workflows
CHI '26· Human-LLM Collaboration +3
- 86%
Situated Imaginaries: Designing AI Futures with Computer Science Teaching Assistants
CHI '26· Human-LLM Collaboration +3
- 83%
From Answer Engines to Learning Partners: A Dual-ZPD Design Framework for AI-Supported Learning
CHI '26· Human-LLM Collaboration +2
- 83%
Designing Scaffolding Cards to Facilitate LLM-Based Socratic Instruction: An Exploratory Study of Response Strategies to Support Learning
CHI '26· Human-LLM Collaboration +2
- 71%
EvaluAId: Human-AI Collaborative Evaluation of Open-Ended Student Essays
CHI '26· Human-LLM Collaboration +3
- 71%
SimStep: Human-in-the-Loop Authoring of Interactive Educational Simulations Through Task-Level Abstractions
CHI '26· Intelligent Tutoring Systems & Learning Analytics +2
- 67%
Classroom Simulacra: Building Contextual Student Generative Agents in Online Education for Learning Behavioral Simulation
CHI '25· Human-LLM Collaboration +1
- 67%
Good Fences Make Good Learning: How Self-Directed Language Learners Navigate LLM Delegation Decisions
CHI '26· Human-LLM Collaboration +1
- 67%
AskNow: An LLM-powered Interactive System for Real-Time Question Answering in Large-Scale Classrooms
CHI '26· Human-LLM Collaboration +1
- 67%
AI meets Mathematics Education: Supporting Instructors in Large Mathematics Classes with Context-Aware AI
CHI '26· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)