Attitudes Surrounding an Imperfect AI Autograder

Explainable AI (XAI)AI-Assisted Decision-Making & AutomationAlgorithmic Fairness & BiasUniversity Professors & ResearchersAI/ML Researchers & Engineers

Title of the Paper

Attitudes Surrounding an Imperfect AI Autograder

Bibliographic Information

  • Subject Area: Human-Computer Interaction and Educational Technology, specifically the application and evaluation of NLP-based automated grading systems in education.
  • Keywords: Human-Computer Interaction, Imperfect AI, AI Acceptance, Folk Theories, Autograder, Computer Science Education, ASAG, Educational Equity, Programming

Research Background and Problem

  • Identified Problems or Challenges:

    1. While AI autograders are widely used in education, research on their interaction with students and students' attitudes toward these systems is relatively scarce.
    2. Students might develop flawed answer-construction strategies when faced with imperfect autograding systems, which could undermine the validity of assessments.
    3. Errors in autograding systems may lead to trust issues, as well as doubts about fairness and educational value.
  • Importance of the Problem:
    As the scale of courses in the education sector grows, the use of AI autograders can significantly reduce labor costs and provide timely feedback. However, the acceptance of these systems by students and instructors may impact the broader adoption of such technologies.

  • Research Motivation and Related Work:
    The authors aim to conduct an in-depth study of students' attitudes toward the autograding systems used in university computer science courses, thereby providing guidance for implementing AI autograders. Prior research on students' attitudes toward autograding systems is limited, particularly regarding the impact of error rates and functional performance on students.

Solution

  • Research Methods:
    The authors employed a mixed-methods approach (including surveys and interviews) to study students' attitudes and interaction experiences. They specifically examined students' answer-construction strategies, perceptions of grader accuracy, views on grading fairness, and their sense of learning value and satisfaction.

  • Key Innovations:

    1. Through quantitative and qualitative analysis of students' attitudes, the authors revealed the impact of students' perceptions of grading errors (especially false negatives and false positives) on their views of accuracy, fairness, and learning value.
    2. The concept of "folk theories" was introduced to explore how students form incorrect or incomplete theories about the mechanisms of autograding systems.
    3. The study proposed several guidelines to integrate imperfect grading systems into classrooms in a more human-centered manner.
  • Implementation Steps and Techniques:
    The study specifically analyzed the error rates of the grading system (15% false positive rate and 10% false negative rate), explored the strategies and misconceptions students developed during their interactions with the system, and formulated preliminary application guidelines:

    • Increase transparency: Help students understand the algorithm's mechanisms and limitations.
    • Use practice-oriented, low-stakes environments (e.g., homework) to help students familiarize themselves with the system.
    • Employ human intervention in high-stakes environments (e.g., exams) to mitigate the impact of grading errors.

Research Outcomes

  • Specific Findings:

    1. Students overestimated the likelihood of the autograder incorrectly marking correct answers as wrong (false negatives), and this overestimation was negatively correlated with their satisfaction and perceptions of fairness.
    2. Students had less awareness of false positives (incorrect answers marked as correct), which could raise concerns about fairness and learning outcomes.
    3. Folk theories and students' perceptions of system errors may lead to suboptimal answer-construction strategies.
  • Advantages Compared to Existing Solutions:
    This study delves into students' subjective experiences, proposes targeted human-centered guidelines, and provides empirical recommendations for the design of future grading systems.

  • Experimental or Evaluation Results:
    Through mixed-methods analysis, the study revealed that students had a limited understanding of the autograder's mechanisms and error rates and exhibited varying degrees of sensitivity to grading errors. Students generally desired greater transparency from the autograder and requested more detailed guidance on dealing with its errors.

  • Limitations and Future Directions:

    • Limitations:
      1. The study's data is context-dependent (university courses) and may not generalize to all autograding systems or educational scenarios.
      2. Self-reported methods might introduce biases, such as recall bias.
    • Future Directions:
      1. Explore unbiased methods to evaluate students' attitudes and cognitive misconceptions.
      2. Develop strategies to enhance the functionality and performance of AI autograders across different domains.
      3. Investigate optimal solutions to reduce the impact of false negatives and false positives on educational equity and learning value.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/47629/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3411764.3445424
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Explainable AI (XAI), AI-Assisted Decision-Making & Automation, Algorithmic Fairness & Bias
work
Professions
University Professors & Researchers, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
8 related papers