AI meets Mathematics Education: Supporting Instructors in Large Mathematics Classes with Context-Aware AI
Authors
Paper Title
AI meets Mathematics Education: Supporting Instructors in Large Mathematics Classes with Context-Aware AI
Publication Info
- Topic area: Application of generative AI in mathematics education for instructional support.
- Keywords: Generative AI, mathematics education, large-enrollment courses, human-centered design, instructional support, pedagogical alignment, AI-assisted learning, hybrid human–AI workflows, course-specific models, automated evaluation.
Background and Problem
- Problem / challenge: Large-enrollment courses face challenges in providing timely and scalable instructional support, especially during peak periods like exams. Existing generative AI systems often lack reliability, pedagogical alignment, and the ability to emulate instructional styles.
- Significance: Addressing this issue can reduce instructor workload, improve student learning experiences, and enhance scalability in education.
- Motivation and related work: Prior research on intelligent tutoring systems and large language models (LLMs) has explored adaptive feedback, personalized learning, and automated evaluation. However, these systems often fail to address pedagogical alignment and nuanced instructional needs. This paper builds on these foundations by focusing on human-centered design and real-world deployment in a university-level Calculus I course.
Solution
- Proposed approach: Development of a lightweight, course-specific generative AI model to provide accurate, context-aware responses to student questions on a course portal.
- Novelty:
- Creation of a course-specific dataset with 2,738 student–instructor Q&A pairs augmented with lecture notes and exercises.
- Fine-tuning of a lightweight language model to emulate the instructional style of course staff.
- Deployment and evaluation of the model in a real-world setting, incorporating instructor oversight and hybrid workflows.
- Analysis of student and instructor perceptions, emphasizing pedagogical alignment and trust.
- Procedure and key techniques:
- Data collection from a Calculus I course portal, including historical Q&A pairs and lecture notes.
- Fine-tuning of selected models using augmented Q&A pairs to enhance scaffolding and instructional alignment.
- Evaluation by a panel of five expert instructors based on correctness, relevance, completeness, and alignment with student intent.
- Deployment on the course portal with instructor oversight, logging actions like endorsements, edits, and deletions.
- Post-deployment student surveys to assess perceptions of model responses.
Results
- Concrete findings:
- The model achieved 75.3% accuracy on a benchmark of 150 representative questions.
- Responses were complete in 64% of cases and relevant in 76.6% of cases.
- In 36% of cases, model responses were rated equal to or better than instructor answers.
- Advantage over baselines:
- Fine-tuned models outperformed base models in pedagogical alignment and accuracy, particularly on medium-difficulty questions.
- The system was 56× more cost-effective than GPT-5 and 3.5× cheaper than DeepSeek.
- Experiments / evaluation:
- Human evaluation by five expert instructors using annotated benchmarks.
- Automated evaluation using GPT-o4-mini for correctness and relevance.
- Post-deployment analysis of instructor actions and student survey responses (N = 105).
- Limitations and future work:
- Limited generalizability due to focus on a single course context.
- Constraints of lightweight models compared to larger architectures.
- Need for advanced automated evaluation methods to reduce reliance on manual annotation.
- Exploration of adaptive response strategies for personalized learning.
Summary
This study presents a lightweight, course-specific generative AI model tailored for instructional support in a large-enrollment Calculus I course. Fine-tuned on 2,738 student–instructor interactions, the model achieved high accuracy and pedagogical alignment, particularly for medium-difficulty questions. Deployment results showed that instructor oversight mitigated risks and increased student trust, highlighting the importance of hybrid human–AI workflows. Students valued the model’s alignment with course materials and immediate availability, though preferences varied regarding response detail. Future work should focus on enhancing pedagogical appropriateness, scalability, and personalization to further integrate AI into STEM education responsibly.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
Classroom Simulacra: Building Contextual Student Generative Agents in Online Education for Learning Behavioral Simulation
CHI '25· Human-LLM Collaboration +1
- 100%
Good Fences Make Good Learning: How Self-Directed Language Learners Navigate LLM Delegation Decisions
CHI '26· Human-LLM Collaboration +1
- 100%
AskNow: An LLM-powered Interactive System for Real-Time Question Answering in Large-Scale Classrooms
CHI '26· Human-LLM Collaboration +1
- 100%
Can an AI Partner Empower Learners to Ask Critical Questions?
IUI '25· Human-LLM Collaboration +1
- 80%
An Interaction Design for Machine Teaching to Develop AI Tutors
CHI '20· Human-LLM Collaboration +2
- 80%
Charting the Future of AI in Project-Based Learning: A Co-Design Exploration with Students
CHI '24· Human-LLM Collaboration +1
- 80%
Unlocking Scientific Concepts: How Effective Are LLM-Generated Analogies for Student Understanding and Classroom Practice?
CHI '25· Human-LLM Collaboration +1
- 80%
Supporting Learners' Use of Imperfect Generative Pedagogical Chatbots: The Role of Chatbot Response Uncertainty and Reduced Verbosity
CHI '26· Conversational Chatbots +2
- 80%
An Empirical Study to Understand How Students Use ChatGPT for Writing Essays
CHI '26· Human-LLM Collaboration +2
- 80%
Who You Explain To Matters: Learning by Explaining to Conversational Agents with Different Pedagogical Roles
CHI '26· Intelligent Tutoring Systems & Learning Analytics +2
Based on Jaccard similarity of research subtopics & professions (≥60%)