Attitudes Surrounding an Imperfect AI Autograder
Authors
Title of the Paper
Attitudes Surrounding an Imperfect AI Autograder
Bibliographic Information
- Subject Area: Human-Computer Interaction and Educational Technology, specifically the application and evaluation of NLP-based automated grading systems in education.
- Keywords: Human-Computer Interaction, Imperfect AI, AI Acceptance, Folk Theories, Autograder, Computer Science Education, ASAG, Educational Equity, Programming
Research Background and Problem
-
Identified Problems or Challenges:
- While AI autograders are widely used in education, research on their interaction with students and students' attitudes toward these systems is relatively scarce.
- Students might develop flawed answer-construction strategies when faced with imperfect autograding systems, which could undermine the validity of assessments.
- Errors in autograding systems may lead to trust issues, as well as doubts about fairness and educational value.
-
Importance of the Problem:
As the scale of courses in the education sector grows, the use of AI autograders can significantly reduce labor costs and provide timely feedback. However, the acceptance of these systems by students and instructors may impact the broader adoption of such technologies. -
Research Motivation and Related Work:
The authors aim to conduct an in-depth study of students' attitudes toward the autograding systems used in university computer science courses, thereby providing guidance for implementing AI autograders. Prior research on students' attitudes toward autograding systems is limited, particularly regarding the impact of error rates and functional performance on students.
Solution
-
Research Methods:
The authors employed a mixed-methods approach (including surveys and interviews) to study students' attitudes and interaction experiences. They specifically examined students' answer-construction strategies, perceptions of grader accuracy, views on grading fairness, and their sense of learning value and satisfaction. -
Key Innovations:
- Through quantitative and qualitative analysis of students' attitudes, the authors revealed the impact of students' perceptions of grading errors (especially false negatives and false positives) on their views of accuracy, fairness, and learning value.
- The concept of "folk theories" was introduced to explore how students form incorrect or incomplete theories about the mechanisms of autograding systems.
- The study proposed several guidelines to integrate imperfect grading systems into classrooms in a more human-centered manner.
-
Implementation Steps and Techniques:
The study specifically analyzed the error rates of the grading system (15% false positive rate and 10% false negative rate), explored the strategies and misconceptions students developed during their interactions with the system, and formulated preliminary application guidelines:- Increase transparency: Help students understand the algorithm's mechanisms and limitations.
- Use practice-oriented, low-stakes environments (e.g., homework) to help students familiarize themselves with the system.
- Employ human intervention in high-stakes environments (e.g., exams) to mitigate the impact of grading errors.
Research Outcomes
-
Specific Findings:
- Students overestimated the likelihood of the autograder incorrectly marking correct answers as wrong (false negatives), and this overestimation was negatively correlated with their satisfaction and perceptions of fairness.
- Students had less awareness of false positives (incorrect answers marked as correct), which could raise concerns about fairness and learning outcomes.
- Folk theories and students' perceptions of system errors may lead to suboptimal answer-construction strategies.
-
Advantages Compared to Existing Solutions:
This study delves into students' subjective experiences, proposes targeted human-centered guidelines, and provides empirical recommendations for the design of future grading systems. -
Experimental or Evaluation Results:
Through mixed-methods analysis, the study revealed that students had a limited understanding of the autograder's mechanisms and error rates and exhibited varying degrees of sensitivity to grading errors. Students generally desired greater transparency from the autograder and requested more detailed guidance on dealing with its errors. -
Limitations and Future Directions:
- Limitations:
- The study's data is context-dependent (university courses) and may not generalize to all autograding systems or educational scenarios.
- Self-reported methods might introduce biases, such as recall bias.
- Future Directions:
- Explore unbiased methods to evaluate students' attitudes and cognitive misconceptions.
- Develop strategies to enhance the functionality and performance of AI autograders across different domains.
- Investigate optimal solutions to reduce the impact of false negatives and false positives on educational equity and learning value.
- Limitations:
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How do students perceive the fairness and accuracy of AI auto-grading systems with error rates?Category: Educational Algorithm Fairness, Learning Opportunity, and Marginalized Student SupportSimilar questionsarrow_forward
- How do error rates (e.g., false negatives and false positives) affect students' satisfaction and perceived learning value of AI auto-grading systems?Category: Educational Algorithm Fairness, Learning Opportunity, and Marginalized Student SupportSimilar questionsarrow_forward
- Will students develop incorrect answer strategies (folk theories) through interaction with AI auto-grading systems?Category: Educational Algorithm Fairness, Learning Opportunity, and Marginalized Student SupportSimilar questionsarrow_forward
Practical Problems
1- Students' distrust of AI grading systems may affect learning outcomes and educational equity.Category: Educational Algorithm Fairness, Learning Opportunity, and Marginalized Student SupportSimilar questionsarrow_forward
- 67%
Deep Learning for Understanding the Human
CHI '18· Human Pose & Activity Recognition +2
- 67%
AI-Moderated Decision-Making: Capturing and Balancing Anchoring Bias in Sequential Decision Tasks
CHI '22· Explainable AI (XAI) +2
- 67%
On Selective, Mutable and Dialogic XAI: a Review of What Users Say about Different Types of Interactive Explanations
CHI '23· Explainable AI (XAI) +1
- 67%
EXMOS: Explanatory Model Steering through Multifaceted Explanations and Data Configurations
CHI '24· Explainable AI (XAI) +1
- 67%
Towards Estimating Missing Emotion Self-reports Leveraging User Similarity: A Multi-task Learning Approach
CHI '24· Explainable AI (XAI) +1
- 67%
Evaluating the Impact of AI-Generated Visual Explanations on Decision-Making for Image Matching
IUI '25· Explainable AI (XAI) +1
- 60%
RetroLens: A Human-AI Collaborative System for Multi-step Retrosynthetic Route Planning
CHI '23· AI-Assisted Decision-Making & Automation
- 60%
Designing Interactive Explainable AI Tools for Algorithmic Literacy and Transparency
DIS '24· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)