Hybrid LLM-Embedded Dialogue Agents for Learner Reflection: Designing Responsive and Theory-Driven Interactions
Authors
Paper Title
Hybrid LLM-Embedded Dialogue Agents for Learner Reflection: Designing Responsive and Theory-Driven Interactions
Publication Info
- Topic area: Hybrid dialogue systems for learner reflection in educational contexts.
- Keywords: Hybrid dialogue systems, Large Language Models, self-regulated learning, learner reflection, culturally responsive computing, open-ended learning, rule-based systems, educational technology, conversational agents, human-computer interaction.
Background and Problem
- Problem / challenge: Traditional rule-based dialogue systems lack flexibility in responding to open-ended learner inputs, while LLMs, though flexible, are not aligned with established learning theories, leading to potential misalignment with pedagogical goals.
- Significance: Supporting learner reflection is critical in open-ended educational environments, as it fosters self-regulated learning (SRL) and deeper engagement with learning tasks.
- Motivation and related work: Prior work has demonstrated the effectiveness of rule-based systems in scaffolding learning and the potential of LLMs for dynamic interactions. However, there is limited research on combining these approaches to balance theoretical grounding with responsiveness. This paper addresses this gap by embedding LLMs within a rule-based framework.
Solution
- Proposed approach: A hybrid dialogue system that integrates a rule-based finite state machine with LLMs to scaffold learner reflection in open-ended learning environments.
- Novelty:
- Design of a hybrid architecture combining rule-based scaffolding with LLM responsiveness.
- A two-stage LLM integration method for relevance checking and context-sensitive follow-up generation.
- Application of the system in a culturally responsive robotics summer camp for middle-school learners.
- Empirical insights into the system’s effectiveness and areas for improvement.
- Procedure and key techniques:
- Rule-based component: Implements structured prompts aligned with the SRL model’s "react and reflect" phase, guiding learners through goal-setting, planning, and reflection.
- LLM-embedded component: Operates in two stages:
- Relevance Check: Determines if learner responses are sufficiently reflective using few-shot examples.
- Contextual Generation: Produces targeted follow-up prompts for deeper reflection when responses are insufficient.
- Deployment: The system was tested in a two-week robotics summer camp, where learners interacted with the system through a chat interface.
Results
- Concrete findings:
- Average session: 39.66 dialogue turns, 6 minutes duration, 3.89 words per learner turn (open-ended).
- LLM was triggered 2.66 times per session on average, with follow-up prompts increasing response length by ~1.75×.
- 9 out of 24 LLM prompts elicited elaborated reflections.
- Advantage over baselines:
- Combines the structured scaffolding of rule-based systems with the generative flexibility of LLMs, enabling context-sensitive and theory-aligned interactions.
- Encouraged richer reflections on goals and activities compared to static rule-based systems.
- Experiments / evaluation:
- Conducted with 9 middle-school learners (ages 9–13) in a culturally responsive robotics camp.
- Mixed-method analysis of dialogue interactions and interviews to assess reflection, engagement, and system performance.
- Metrics: word count, turn count, LLM trigger frequency, and qualitative coding of learner responses.
- Limitations and future work:
- Contextual misalignment: LLM struggled with aesthetic design contexts (e.g., hair and accessories).
- Affective misalignment: Failed to recognize disengagement cues, leading to frustration in some learners.
- Missed opportunities for deeper reflection due to overly permissive relevance checks.
- Small sample size (N = 9) limits generalizability.
- Future work: Incorporate domain-specific knowledge (e.g., Retrieval Augmented Generation), improve emotional attunement, adapt to learner preferences, and explore co-creation paradigms.
Summary
This paper presents a hybrid dialogue system that integrates rule-based scaffolding with LLM responsiveness to support learner reflection in open-ended educational contexts. Tested in a culturally responsive robotics summer camp, the system effectively elicited reflections on goals and activities but faced challenges with contextual and affective alignment. The findings highlight the potential of hybrid systems to balance pedagogical grounding with flexibility, offering design insights for future applications in education and beyond.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
Exploring Learners' Expectations and Engagement When Collaborating with Constructively Controversial Peer Agents
CHI '26· Human-LLM Collaboration +2
- 71%
Designing AI Peers for Collaborative Mathematical Problem Solving with Middle School Students: A Participatory Design Study
CHI '26· Intelligent Tutoring Systems & Learning Analytics +2
- 71%
Unveiling the Capabilities of Large Language Models in Simulating Student Behavioral Dynamics and Supporting Peer Feedback to Augment Task Performance
CHI '26· Human-LLM Collaboration +2
- 67%
Adaptive Empathy Learning Support in Peer Review Scenarios
CHI '22· Intelligent Tutoring Systems & Learning Analytics +1
- 67%
Classroom Simulacra: Building Contextual Student Generative Agents in Online Education for Learning Behavioral Simulation
CHI '25· Human-LLM Collaboration +1
- 67%
Good Fences Make Good Learning: How Self-Directed Language Learners Navigate LLM Delegation Decisions
CHI '26· Human-LLM Collaboration +1
- 67%
AskNow: An LLM-powered Interactive System for Real-Time Question Answering in Large-Scale Classrooms
CHI '26· Human-LLM Collaboration +1
- 67%
AI meets Mathematics Education: Supporting Instructors in Large Mathematics Classes with Context-Aware AI
CHI '26· Human-LLM Collaboration +1
- 67%
Can an AI Partner Empower Learners to Ask Critical Questions?
IUI '25· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)