Playing Dumb to Get Smart: Creating and Evaluating an LLM-based Teachable Agent within University Computer Science Classes
Authors
Research Background and Issues
-
What problems or challenges did the authors identify?
When students learn programming through the "Learning by Teaching" (LBT) method, it can effectively deepen their understanding. However, traditional human-based LBT is limited in scalability and risks reinforcing misconceptions. Therefore, there is a need for a teachable agent driven by large language models (LLMs) to simulate the role of non-expert learners and expand the scope of LBT practices. -
Why is this issue important?
LBT not only helps students consolidate knowledge but also promotes active learning, identifies knowledge gaps, and improves attitudes toward programming. However, tools for implementing this method often rely on human or even expert resources, making it difficult to scale in larger educational environments. Additionally, while students traditionally use LLMs as knowledge-providing tools, there is a lack of research exploring their potential and long-term impact as "teachable learners." -
Research Motivation and Related Work
Existing studies show that using AI or robots as teaching objects can improve student learning outcomes, and LLMs, with their conversational generation capabilities, clearly have the potential to simulate the role of learners. However, the long-term impact of LLM-based LBT in extended course environments requires further empirical investigation.
Solution
-
What methods or solutions did the authors propose?
This study designed and evaluated an LLM-based teachable agent tool—MatlabTutee—using Chain-of-Thought prompts to guide LLMs in simulating the role of non-expert learners, enhancing students' LBT experiences. It also compared this system with humans pretending to be non-experts and explored its long-term use in real university computer science courses. -
What are the innovative aspects of this solution?
- Using simple Chain-of-Thought prompts to control LLM-generated teaching dialogues without complex knowledge state simulations.
- Incorporating multiple intentional errors and typical non-expert behaviors to maintain consistency in the learner role.
- Directly comparing the system with traditional LLMs and humans pretending to be learners to quantify its performance as a teachable agent and its interaction with students.
-
What are the implementation steps and key technologies used?
- Preliminary Design and Evaluation: Developing initial prompts and observing student interactions with LLMs or human assistants through single-day experiments.
- Chain-of-Thought Optimization: Refining prompts based on feedback to introduce Chain-of-Thought methods and test the system's ability to generate incorrect code and explanations.
- Comprehensive Comparative Experiments: Conducting structured week-long experiments comparing MatlabTutee with similar human and traditional LLM systems, analyzing interaction quality, student behavior, and cognitive load.
- Long-term Usage Studies: Performing two month-long longitudinal experiments (partially controlled and fully uncontrolled) to evaluate students' autonomous adoption of the system in real teaching scenarios and its long-term effects.
Research Outcomes
-
What specific outcomes were achieved?
- MatlabTutee successfully simulated the role of a non-expert learner, with its intelligence level perceived by students as similar to humans pretending to be uninformed experts.
- Interacting with MatlabTutee enabled students to identify knowledge gaps, accurately assess their own abilities, and improve their positive attitude toward computer science.
-
What advantages does it have compared to existing solutions?
- Compared to traditional GPT models, MatlabTutee better facilitates active student engagement in learning rather than simply providing answers from AI.
- Its interaction model resembles human learners, effectively stimulating active knowledge construction and deep reflection during the LBT process.
-
What were the experimental or evaluation results?
- Under controlled conditions, MatlabTutee's LBT experience was similar to that of human interactions, prompting students to engage in active teaching behaviors such as explanation, reasoning, and deep learning. In contrast, traditional GPT models were primarily used for knowledge retrieval, leading to more passive learning behaviors.
- In long-term experiments, while most students recognized the system's potential benefits, they reduced usage frequency due to incorrect feedback and conflicts with course schedules, with average interaction times being relatively short.
-
Limitations and Future Directions
- MatlabTutee exhibited "learning too quickly" during extended interactions, leading to distrust among some students, necessitating further optimization of simulated learning trajectory design.
- Students had polarized reactions to teaching errors, which may impact sustained usage. Future designs could explore more motivational approaches to enhance feedback acceptance.
- Longitudinal experiments revealed low student autonomy in adopting the system, suggesting the need for integration with course objectives or grading mechanisms to encourage broader adoption.
Conclusion and Discussion
- This study demonstrated the potential of Chain-of-Thought prompts in constructing simple yet effective LLM-based teachable agents, while also revealing challenges such as insufficient student intent for natural usage, sensitivity to incorrect feedback, and inconsistent simulated learning trajectories.
- Open questions for future research include balancing transparency with customized learning trajectories, providing constructive feedback that does not undermine motivation, and expanding students' long-term autonomous use of such systems.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can LLMs simulate non-expert learners to improve application of learning-by-teaching (LBT) methods?Category: LLM Learning Scaffolding and Reflection SupportSimilar questionsarrow_forward
- What innovative methods exist for designing LLM-based teachable agents to enhance students' teaching experience?Category: LLM Learning Scaffolding and Reflection SupportSimilar questionsarrow_forward
- What impact will introducing LLM-based LBT tools in programming courses have on students' learning behavior and long-term outcomes?Category: LLM Learning Scaffolding and Reflection SupportSimilar questionsarrow_forward
Practical Problems
1- Students lack scalable and efficient tools to practice learning-by-teaching methods.Category: LLM Learning Scaffolding and Reflection SupportSimilar questionsarrow_forward
- 100%
Exploring the Design Space of Cognitive Engagement Techniques with AI-Generated Code for Enhanced Learning
IUI '25· Human-LLM Collaboration +1
- 80%
DBox: Scaffolding Algorithmic Programming Learning through Learner-LLM Co-Decomposition
CHI '25· Human-LLM Collaboration +2
- 80%
LingoQ: Bridging the Gap between EFL Learning and Work through AI-Generated Work-Related Quizzes
CHI '26· Human-LLM Collaboration +2
- 67%
From Code Generation to Conceptual Learning: Student Use of LLMs in a Web Programming Course
CHI '26· Human-LLM Collaboration +2
- 67%
Exploring the Learnability of Program Synthesizers by Novice Programmers
UIST '22· Human-LLM Collaboration +1
- 60%
CodeAid: Evaluating a Classroom Deployment of an LLM-based Programming Assistant that Balances Student and Educator Needs
CHI '24· Human-LLM Collaboration +1
- 60%
BIDTrainer: An LLMs-driven Education Tool for Enhancing the Understanding and Reasoning in Bio-inspired Design
CHI '24· Human-LLM Collaboration +1
- 60%
Teach AI How to Code: Using Large Language Models as Teachable Agents for Programming Education
CHI '24· Human-LLM Collaboration +1
- 60%
PromptHive: Bringing Subject Matter Experts Back to the Forefront with Collaborative Prompt Engineering for Educational Content Creation
CHI '25· Human-LLM Collaboration +1
- 60%
How Humans Communicate Programming Tasks in Natural Language and Implications For End-User Programming with LLMs
CHI '25· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)