Unveiling the Capabilities of Large Language Models in Simulating Student Behavioral Dynamics and Supporting Peer Feedback to Augment Task Performance
Paper Title
Unveiling the Capabilities of Large Language Models in Simulating Student Behavioral Dynamics and Supporting Peer Feedback to Augment Task Performance
Publication Info
- Topic area: Application of large language models (LLMs) in student simulation and peer feedback for educational enhancement.
- Keywords: Large language models, student simulation, peer feedback, metacognition, learning science, cognitive performance, virtual peers, educational AI, simulation experiments, behavioral dynamics.
Background and Problem
- Problem / challenge: Existing LLM-based student simulators primarily focus on coarse-grained metrics like question-answering accuracy, neglecting finer-grained behavioral dynamics such as prior knowledge, contextual experience, and sensory actions. Explanations for LLMs’ successes and failures in simulations are limited.
- Significance: Understanding and improving LLMs’ ability to simulate student behaviors can enhance educational research, content design, instructional strategies, and scalable student support systems.
- Motivation and related work: Prior research has explored LLMs in educational contexts, including tutoring, assessment, and collaborative learning. However, most simulators operate at a coarse granularity, missing nuanced learning trajectories. This paper addresses these gaps by conducting fine-grained simulation experiments and proposing new methodologies.
Solution
- Proposed approach: Meta-Cognitive Refinement (MCR), a dual-layer refinement process integrating metacognitive reasoning at both the student and model levels to enhance simulation fidelity.
- Novelty:
- Conducted large-scale, fine-grained simulation experiments to systematically investigate LLMs’ capabilities and limitations in modeling student behaviors.
- Introduced MCR, leveraging insights from simulations and grounding mechanisms in metacognitive regulation and learning science theories.
- Demonstrated LLM-powered virtual peers’ effectiveness in providing peer-pressure feedback to improve cognitive task performance.
- Procedure and key techniques:
- Conducted six simulation experiments using diverse datasets to analyze LLMs’ ability to model student demographics, learning history, course material comprehension, contextual knowledge, prior knowledge, engagement, and behavioral correlations.
- Developed MCR to refine simulations through dual-layer reasoning: student-level reasoning interprets task demands and behavioral history, while model-level reasoning ensures internal consistency and alignment.
- Evaluated MCR against baseline models (Standard, CoT, PoT, Self-Refine) using metrics like Mean Absolute Error (MAE) and correlation coefficients.
- Conducted a user study with virtual peers providing feedback in math tasks to assess their impact on response time and accuracy.
Results
- Concrete findings:
- Strong alignment between simulated and real student behaviors for demographic factors (e.g., note-taking habits, r = 0.94) and learning history (r = 0.69).
- Incorporating contextual knowledge and engagement significantly improved simulation accuracy (e.g., Tcontext − pre − engage, r = 0.545).
- MCR achieved the lowest simulation error (MAE = 3.225) and strongest correlation with real response time sequences (r = 0.375, p =.007).
- Virtual peers accelerated response times significantly compared to control and random peer groups (β = 0.182, p <.001).
- Advantage over baselines:
- MCR outperformed Standard, CoT, PoT, and Self-Refine models in simulation fidelity, reducing MAE and improving correlation with real student data.
- Virtual peers using MCR feedback achieved greater performance improvements than real peers and random peers.
- Experiments / evaluation:
- Six simulation experiments tested LLMs’ ability to model diverse learning scenarios using datasets like Nevriye et al. (N = 145), Open University Learning Analytics Dataset (N = 4524), and EduAgent (N = 311).
- User study with N = 188 participants compared cognitive task performance across Control, Real Peer, Virtual Peer, and Random Peer groups.
- Limitations and future work:
- LLM sensitivity to prompt design and token size constraints may limit broader applications.
- Potential biases in LLMs’ internal knowledge base could affect simulation accuracy.
- Future work should explore richer student-specific data, complex domains, and diverse virtual peer interactions.
Summary
This paper systematically investigates the capabilities and limitations of LLMs in simulating student behaviors through fine-grained experiments and introduces Meta-Cognitive Refinement (MCR) to enhance simulation fidelity. MCR leverages insights from learning science and metacognitive theories to refine simulations at both the student and model levels. Results demonstrate that MCR-powered virtual peers effectively accelerate cognitive task performance without compromising accuracy, highlighting LLMs’ potential as scalable, interpretable tools for educational research and adaptive learning systems. Future work should address biases, expand domains, and explore richer personalization.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
Exploring Learners' Expectations and Engagement When Collaborating with Constructively Controversial Peer Agents
CHI '26· Human-LLM Collaboration +2
- 71%
AutoPBL: An LLM-powered Platform to Guide and Support Individual Learners Through Self Project-based Learning
CHI '25· Human-LLM Collaboration +2
- 71%
Hybrid LLM-Embedded Dialogue Agents for Learner Reflection: Designing Responsive and Theory-Driven Interactions
CHI '26· Human-LLM Collaboration +2
- 71%
Designing AI Peers for Collaborative Mathematical Problem Solving with Middle School Students: A Participatory Design Study
CHI '26· Intelligent Tutoring Systems & Learning Analytics +2
- 71%
Co-Designing with Algorithms: Unpacking the Complex Role of GenAI in Interactive System Design Education
DIS '25· Generative AI (Text, Image, Music, Video) +2
- 67%
Adaptive Empathy Learning Support in Peer Review Scenarios
CHI '22· Intelligent Tutoring Systems & Learning Analytics +1
- 67%
Classroom Simulacra: Building Contextual Student Generative Agents in Online Education for Learning Behavioral Simulation
CHI '25· Human-LLM Collaboration +1
- 67%
Good Fences Make Good Learning: How Self-Directed Language Learners Navigate LLM Delegation Decisions
CHI '26· Human-LLM Collaboration +1
- 67%
AskNow: An LLM-powered Interactive System for Real-Time Question Answering in Large-Scale Classrooms
CHI '26· Human-LLM Collaboration +1
- 67%
AI meets Mathematics Education: Supporting Instructors in Large Mathematics Classes with Context-Aware AI
CHI '26· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)