AI-Mediated Feedback Improves Student Revisions: A Randomized Trial with FeedbackWriter in a Large Undergraduate Course
Authors
Paper Title
AI-Mediated Feedback Improves Student Revisions: A Randomized Trial with FeedbackWriter in a Large Undergraduate Course
Publication Info
- Topic area: AI-mediated feedback in educational settings
- Keywords: AI-mediated feedback, FeedbackWriter, large language models, knowledge-intensive essays, human-AI collaboration, revision quality, teaching assistants, rubric-driven feedback, student learning, randomized controlled trial
Background and Problem
- Problem / challenge: Existing research on AI-generated feedback shows mixed results, with limited exploration of AI-mediated feedback where human evaluators collaborate with AI. Prior studies often focus on short-answer questions or evaluate feedback quality without measuring its impact on student outcomes.
- Significance: Understanding AI-mediated feedback's effectiveness in improving student revisions and learning can inform scalable, high-quality feedback provision in large courses, addressing challenges like instructor workload and feedback consistency.
- Motivation and related work: Prior work has shown that AI-generated feedback can be useful but often lacks accuracy and alignment with teaching goals. Human-AI collaboration has shown promise in other domains, but its application to knowledge-intensive essay feedback remains underexplored. This study builds on these gaps by investigating AI-mediated feedback in a real-world classroom setting.
Solution
- Proposed approach: FeedbackWriter, a system that provides AI-generated feedback suggestions to teaching assistants (TAs) while allowing them to review, edit, or dismiss the suggestions.
- Novelty:
- Introduction of FeedbackWriter, aligning AI feedback generation with human evaluators’ workflows for knowledge-intensive essays.
- First randomized controlled trial comparing AI-mediated feedback to human-only feedback on student revision quality and learning outcomes.
- Iterative rubric refinement to improve AI judgment accuracy and feedback quality.
- Analysis of TA engagement with AI suggestions and their impact on feedback quality and student outcomes.
- Procedure and key techniques:
- FeedbackWriter generates feedback by identifying relevant sentences, judging rubric satisfaction, and drafting feedback messages.
- TAs use an interactive interface to review AI suggestions, modify judgments, and provide feedback aligned with rubrics.
- A randomized trial was conducted in an undergraduate economics course (N=354 students, 11 TAs), with students receiving either AI-mediated or human-only feedback on two writing assignments.
Results
- Concrete findings:
- Students receiving AI-mediated feedback produced higher-quality revisions (Cohen’s d = 0.50), equivalent to moving from the 50th to the 70th percentile.
- AI-mediated feedback exhibited higher actionability (89.3% vs. 75.8%), promotion of independent learning (92.6% vs. 82.3%), and tone/supportiveness (97.4% vs. 80.8%) compared to human-only feedback.
- TAs adopted AI suggestions for 51.3% of feedback comments and corrected AI judgments in 11.3% of cases.
- Advantage over baselines:
- AI-mediated feedback resulted in more comprehensive and rubric-aligned feedback, with higher inter-rater consistency among TAs.
- Students receiving AI-mediated feedback showed greater improvement in revision quality compared to those receiving human-only feedback.
- Experiments / evaluation:
- Randomized controlled trial with counterbalanced conditions across two assignments.
- Evaluation metrics included revision quality (using an LLM-based rubric scorer), post-test scores, and feedback quality dimensions.
- Data sources included system logs, qualitative analysis of feedback, and TA interviews.
- Limitations and future work:
- Lack of comparison with AI-only feedback.
- Use of AI-based evaluators for essay quality, which may introduce bias.
- Limited exploration of long-term learning outcomes and student perceptions of AI-mediated feedback.
- Future work should investigate sustained practice effects, refine rubrics, and explore real-time adaptation of AI suggestions.
Summary
This study introduces FeedbackWriter, a system that enables TAs to provide AI-mediated feedback on knowledge-intensive essays. A randomized trial in a large undergraduate course demonstrated that AI-mediated feedback improves student revision quality and exhibits desirable feedback properties like actionability and independent learning promotion. TAs actively engaged with AI suggestions, enhancing feedback comprehensiveness and consistency. The findings highlight the potential of human-AI collaboration in scaling high-quality feedback, particularly in structured, rubric-driven assignments. Future research should explore AI-only feedback, long-term learning impacts, and broader adoption across diverse educational contexts.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
An Empirical Study to Understand How Students Use ChatGPT for Writing Essays
CHI '26· Human-LLM Collaboration +2
- 80%
Classroom Simulacra: Building Contextual Student Generative Agents in Online Education for Learning Behavioral Simulation
CHI '25· Human-LLM Collaboration +1
- 80%
Good Fences Make Good Learning: How Self-Directed Language Learners Navigate LLM Delegation Decisions
CHI '26· Human-LLM Collaboration +1
- 80%
AskNow: An LLM-powered Interactive System for Real-Time Question Answering in Large-Scale Classrooms
CHI '26· Human-LLM Collaboration +1
- 80%
AI meets Mathematics Education: Supporting Instructors in Large Mathematics Classes with Context-Aware AI
CHI '26· Human-LLM Collaboration +1
- 80%
Can an AI Partner Empower Learners to Ask Critical Questions?
IUI '25· Human-LLM Collaboration +1
- 67%
An Interaction Design for Machine Teaching to Develop AI Tutors
CHI '20· Human-LLM Collaboration +2
- 67%
Charting the Future of AI in Project-Based Learning: A Co-Design Exploration with Students
CHI '24· Human-LLM Collaboration +1
- 67%
Unlocking Scientific Concepts: How Effective Are LLM-Generated Analogies for Student Understanding and Classroom Practice?
CHI '25· Human-LLM Collaboration +1
- 67%
Supporting Learners' Use of Imperfect Generative Pedagogical Chatbots: The Role of Chatbot Response Uncertainty and Reduced Verbosity
CHI '26· Conversational Chatbots +2
Based on Jaccard similarity of research subtopics & professions (≥60%)