Classroom Simulacra: Building Contextual Student Generative Agents in Online Education for Learning Behavioral Simulation
Authors
Research Background and Issues
- Issues and Challenges: The authors point out that existing student simulation studies and methods generally overlook the modulation effect of course materials on student learning behaviors. The main reasons include:
- Lack of high-quality datasets containing fine-grained course materials and real-time student learning performance.
- Existing simulation models (especially language models) struggle to handle long-context texts such as course materials.
- Significance: Accurately simulating student learning behaviors can support the construction of "digital twin" classrooms, which is crucial for educators to explore new teaching strategies, enhance student learning outcomes, and develop intelligent education systems.
- Research Motivation: Although generative artificial intelligence (e.g., large language models, LLMs) has shown potential in predicting student performance, its effectiveness in realistic simulations has not been systematically validated. Moreover, current research focuses more on predicting student scores rather than capturing the dynamic changes in learning behaviors.
- Related Work:
- Previous methods primarily focus on student knowledge tracing, such as using deep learning and Bayesian models to predict learning behaviors.
- Some studies in educational systems have introduced generative agents as teaching assistants but failed to fully integrate course materials or focus on realistic behavior simulation.
Solution
- Proposed Method: The authors propose a "Transferable Iterative Reflection" (TIR) module that combines large language models (LLMs) with course materials to enhance the simulation of student learning behaviors.
- Innovations: The TIR module compresses knowledge by guiding LLMs through multiple rounds of reflection, enabling it to:
- Iteratively adjust prediction biases.
- Extract transferable reflective information to enhance reasoning capabilities.
- Implementation Steps:
- Data Collection:
- Conducted a 6-week teaching experiment using a self-developed online education system, collecting fine-grained behavioral data and course materials from 60 students.
- The system utilized multimodal sensing technologies (e.g., gaze tracking, emotion detection) to improve data quality.
- Model Framework:
- Proposed three student simulation models: prompt-based LLMs, fine-tuned LLMs, and deep learning-based knowledge tracing models (as baselines).
- The TIR module consists of four stages:
- Initial Prediction: The model makes its first prediction of students' future performance.
- Reflection: The model generates reflections based on discrepancies between the original predictions and labels.
- Testing: The reflective knowledge assists the new model in generating more accurate predictions.
- Iteration: The process is repeated until optimal results are achieved.
- Model Enhancement:
- For prompt-based models, the TIR module enhances the contextual learning efficiency through example reflections.
- For fine-tuned models, the TIR module compresses contextual knowledge, addressing the token limitation issue of LLMs.
- Data Collection:
Research Outcomes
- Specific Results:
- Constructed a 6-week experimental dataset containing high-quality learning behaviors and course materials, supporting finer-grained student simulations.
- The proposed TIR module significantly improved the simulation capabilities of both prompt-based and fine-tuned LLMs, outperforming traditional deep learning-based models.
- Advantages Comparison:
- Compared to deep learning baseline models:
- TIR-enhanced fine-tuned language models (e.g., BertKT) outperformed the best deep learning model (SimpleKT) in both accuracy and F1 scores.
- Compared to the original large language models:
- Prompt-based models achieved significant improvements in data utilization efficiency through the TIR module, performing well even with minimal training data (only 4 training examples).
- The TIR module also significantly enhanced the performance of smaller LLMs (e.g., GPT-4o Mini), enabling them to achieve simulation accuracy comparable to or even surpassing larger models (GPT-4o).
- Compared to deep learning baseline models:
- Experimental Results:
- Validated the model's superiority on a public dataset (EduAgent), where BertKT+TIR achieved an accuracy of 70.12% (approximately 3% higher than the best deep learning model).
- On newly collected data, the TIR module demonstrated better capture of individual differences, course difficulty, and question relevance.
- Iterative reflection captured fine-grained dimensions: individual level, course level, question level, and skill change pathways.
- For instance, the TIR-enhanced model improved the correlation coefficient between simulated and real student average accuracy from 0.02 to 0.42.
- Simulated student groups also exhibited interaction patterns more consistent with real students.
- Limitations and Future Directions:
- Sample Diversity: The current experiment focuses on primary and secondary school students. Future research could expand to more diverse populations (e.g., university students, adult learners).
- Behavior Types: The study currently simulates only students' answer correctness, without addressing more complex learning behaviors (e.g., learning styles, cognitive processes). Future studies could incorporate additional dimensions of behavior simulation.
- Model Application Scenarios: Although TIR improves student simulation performance, further testing in real educational scenarios is needed to fully validate its value in teaching improvement.
Conclusion
This study proposes an innovative solution for improving learning behavior simulation in online education through a generative student agent based on transferable iterative reflection. It not only advances simulation technology but also demonstrates its broad potential in future educational systems, particularly in personalized learning path design and teaching strategy optimization.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How do course materials regulate students' learning behavior, and how can language models simulate this regulatory effect?Category: Learning Support Needs and Educational Interaction Pain PointsSimilar questionsarrow_forward
- How can limitations of language models in processing long texts (e.g., course materials) be overcome?Category: Learning Support Needs and Educational Interaction Pain PointsSimilar questionsarrow_forward
- How does the Transferable Iterative Reflection (TIR) module improve simulation of students' learning behavior?Category: Learning Support Needs and Educational Interaction Pain PointsSimilar questionsarrow_forward
Practical Problems
1- Precise simulation of students' learning behavior in online education is difficult, hindering optimization of teaching strategies.Category: Learning Support Needs and Educational Interaction Pain PointsSimilar questionsarrow_forward
- 100%
Good Fences Make Good Learning: How Self-Directed Language Learners Navigate LLM Delegation Decisions
CHI '26· Human-LLM Collaboration +1
- 100%
AskNow: An LLM-powered Interactive System for Real-Time Question Answering in Large-Scale Classrooms
CHI '26· Human-LLM Collaboration +1
- 100%
AI meets Mathematics Education: Supporting Instructors in Large Mathematics Classes with Context-Aware AI
CHI '26· Human-LLM Collaboration +1
- 100%
Can an AI Partner Empower Learners to Ask Critical Questions?
IUI '25· Human-LLM Collaboration +1
- 80%
An Interaction Design for Machine Teaching to Develop AI Tutors
CHI '20· Human-LLM Collaboration +2
- 80%
Charting the Future of AI in Project-Based Learning: A Co-Design Exploration with Students
CHI '24· Human-LLM Collaboration +1
- 80%
Unlocking Scientific Concepts: How Effective Are LLM-Generated Analogies for Student Understanding and Classroom Practice?
CHI '25· Human-LLM Collaboration +1
- 80%
Supporting Learners' Use of Imperfect Generative Pedagogical Chatbots: The Role of Chatbot Response Uncertainty and Reduced Verbosity
CHI '26· Conversational Chatbots +2
- 80%
An Empirical Study to Understand How Students Use ChatGPT for Writing Essays
CHI '26· Human-LLM Collaboration +2
- 80%
Who You Explain To Matters: Learning by Explaining to Conversational Agents with Different Pedagogical Roles
CHI '26· Intelligent Tutoring Systems & Learning Analytics +2
Based on Jaccard similarity of research subtopics & professions (≥60%)