AI meets Mathematics Education: Supporting Instructors in Large Mathematics Classes with Context-Aware AI

Human-LLM CollaborationIntelligent Tutoring Systems & Learning AnalyticsUniversity Professors & ResearchersOnline Course Designers

Paper Title

AI meets Mathematics Education: Supporting Instructors in Large Mathematics Classes with Context-Aware AI

Publication Info

  • Topic area: Application of generative AI in mathematics education for instructional support.
  • Keywords: Generative AI, mathematics education, large-enrollment courses, human-centered design, instructional support, pedagogical alignment, AI-assisted learning, hybrid human–AI workflows, course-specific models, automated evaluation.

Background and Problem

  • Problem / challenge: Large-enrollment courses face challenges in providing timely and scalable instructional support, especially during peak periods like exams. Existing generative AI systems often lack reliability, pedagogical alignment, and the ability to emulate instructional styles.
  • Significance: Addressing this issue can reduce instructor workload, improve student learning experiences, and enhance scalability in education.
  • Motivation and related work: Prior research on intelligent tutoring systems and large language models (LLMs) has explored adaptive feedback, personalized learning, and automated evaluation. However, these systems often fail to address pedagogical alignment and nuanced instructional needs. This paper builds on these foundations by focusing on human-centered design and real-world deployment in a university-level Calculus I course.

Solution

  • Proposed approach: Development of a lightweight, course-specific generative AI model to provide accurate, context-aware responses to student questions on a course portal.
  • Novelty:
    1. Creation of a course-specific dataset with 2,738 student–instructor Q&A pairs augmented with lecture notes and exercises.
    2. Fine-tuning of a lightweight language model to emulate the instructional style of course staff.
    3. Deployment and evaluation of the model in a real-world setting, incorporating instructor oversight and hybrid workflows.
    4. Analysis of student and instructor perceptions, emphasizing pedagogical alignment and trust.
  • Procedure and key techniques:
    • Data collection from a Calculus I course portal, including historical Q&A pairs and lecture notes.
    • Fine-tuning of selected models using augmented Q&A pairs to enhance scaffolding and instructional alignment.
    • Evaluation by a panel of five expert instructors based on correctness, relevance, completeness, and alignment with student intent.
    • Deployment on the course portal with instructor oversight, logging actions like endorsements, edits, and deletions.
    • Post-deployment student surveys to assess perceptions of model responses.

Results

  • Concrete findings:
    • The model achieved 75.3% accuracy on a benchmark of 150 representative questions.
    • Responses were complete in 64% of cases and relevant in 76.6% of cases.
    • In 36% of cases, model responses were rated equal to or better than instructor answers.
  • Advantage over baselines:
    • Fine-tuned models outperformed base models in pedagogical alignment and accuracy, particularly on medium-difficulty questions.
    • The system was 56× more cost-effective than GPT-5 and 3.5× cheaper than DeepSeek.
  • Experiments / evaluation:
    • Human evaluation by five expert instructors using annotated benchmarks.
    • Automated evaluation using GPT-o4-mini for correctness and relevance.
    • Post-deployment analysis of instructor actions and student survey responses (N = 105).
  • Limitations and future work:
    • Limited generalizability due to focus on a single course context.
    • Constraints of lightweight models compared to larger architectures.
    • Need for advanced automated evaluation methods to reduce reliance on manual annotation.
    • Exploration of adaptive response strategies for personalized learning.

Summary

This study presents a lightweight, course-specific generative AI model tailored for instructional support in a large-enrollment Calculus I course. Fine-tuned on 2,738 student–instructor interactions, the model achieved high accuracy and pedagogical alignment, particularly for medium-difficulty questions. Deployment results showed that instructor oversight mitigated risks and increased student trust, highlighting the importance of hybrid human–AI workflows. Students valued the model’s alignment with course materials and immediate availability, though preferences varied regarding response detail. Future work should focus on enhancing pedagogical appropriateness, scalability, and personalization to further integrate AI into STEM education responsibly.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223455/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791236
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Human-LLM Collaboration, Intelligent Tutoring Systems & Learning Analytics
work
Professions
University Professors & Researchers, Online Course Designers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers