SimStep: Human-in-the-Loop Authoring of Interactive Educational Simulations Through Task-Level Abstractions
Authors
Paper Title
SimStep: Human-in-the-Loop Authoring of Interactive Educational Simulations Through Task-Level Abstractions
Publication Info
- Topic area: AI-assisted educational simulation authoring
- Keywords: generative AI, human-in-the-loop, task-level abstractions, educational simulations, distributed cognition, Chain-of-Abstractions, pedagogical reasoning, interactive learning, simulation design, debugging
Background and Problem
- Problem / challenge: Generative AI tools for creating educational simulations lack programming affordances such as traceability, refinement, and debugging, making it difficult for educators to align simulations with learning goals or verify their fidelity to intended concepts.
- Significance: Addressing these limitations would enable educators to create more effective and tailored interactive learning experiences, enhancing STEM education and discovery learning.
- Motivation and related work: Prior work in Intelligent Tutoring Systems (ITS) and prompt-based programming has explored domain-aligned representations and debugging interfaces but often struggles with high cognitive load, limited expressiveness, or insufficient error correction. SimStep builds on these approaches by introducing structured task-level abstractions to reduce complexity and improve control.
Solution
- Proposed approach: SimStep, a human-in-the-loop authoring environment that structures simulation design through a Chain-of-Abstractions (CoA) framework, enabling educators to specify, inspect, and refine simulations step-by-step.
- Novelty:
- Multi-layer task-level abstractions (Concept Graph, Scenario Graph, Learning Goal Graph, UI Graph) aligned with pedagogical reasoning.
- Inverse correction process for revising hidden model assumptions at abstraction levels rather than code.
- End-to-end traceability of concepts through abstraction stages, preserving fidelity and enabling iterative refinement.
- Procedure and key techniques:
- Educators input learning content, which is transformed into a Concept Graph.
- Scenario selection contextualizes the Concept Graph into a Scenario Graph.
- Learning goals filter the Scenario Graph into a Learning Goal Graph.
- The Learning Goal Graph generates a UI Graph, which is translated into simulation code.
- Errors are corrected through guided testing, automated repair, or manual debugging using abstraction widgets.
Results
- Concrete findings:
- Average usability score: 4.66 (PSSUQ scale, 1–6).
- Average abstraction generation time: Concept Map (3.54s), Scenario Map (3.45s), Learning Goal Map (3.27s), UI Map & Simulation Code (36.27s).
- 48.5% of simulations worked without intervention; 73.5% of broken simulations were fixed using inverse correction widgets.
- Advantage over baselines:
- SimStep offers multi-layer pre-code abstractions, cross-layer inverse correction, and end-to-end traceability, surpassing other systems in repair scope and pedagogical alignment.
- Experiments / evaluation:
- User study with 11 STEM educators (average teaching experience: 10.18 years) evaluated usability, task load, and cognitive dimensions of abstractions.
- Fidelity evaluation with 66 curated simulations rated abstractions on completeness/coherence (highest: Concept Map, μ = 2.61), inclusion consistency (μ = 2.53), and meaning drift (μ = 2.57) on a 3-point scale.
- Limitations and future work:
- Challenges in abstraction customization for non-standard workflows.
- Cognitive load introduced by layered representations.
- Sensitivity of transformations to minor input changes.
- Lack of demographic data for abstraction fidelity evaluators.
Summary
SimStep introduces a Chain-of-Abstractions framework to enable educators to author interactive STEM simulations through structured, human-in-the-loop workflows. By decomposing simulation design into task-level abstractions, SimStep provides traceability, refinement, and error correction, addressing limitations of generative AI tools. User studies and fidelity evaluations demonstrate its effectiveness in reducing authoring complexity, supporting pedagogical alignment, and enabling iterative debugging. Future work will explore adaptive interfaces, abstraction customization, and robustness to input perturbations.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 86%
Codesigning Ripplet: an LLM-Assisted Assessment Authoring System Grounded in a Conceptual Model of Teachers’ Workflows
CHI '26· Human-LLM Collaboration +3
- 71%
Thinking in Graphs with CoMAP: A Shared Visual Workspace for Designing Project--Based Learning
CHI '26· Collaborative Learning & Peer Teaching +2
- 71%
Open-ended Structured Question Assessment with Human-LLM Collaboration
CHI '26· Human-LLM Collaboration +3
- 71%
Designing Looms as Kits for Collaborative Assembly
UIST '25· Makerspace Culture +2
- 63%
"Listen to the Teachers": Research-Based Personas for Translating Classroom Realities into Actionable HCI Design
CHI '26· User Research Methods (Interviews, Surveys, Observation) +3
- 63%
Situated Imaginaries: Designing AI Futures with Computer Science Teaching Assistants
CHI '26· Human-LLM Collaboration +3
- 63%
OpenCD: Empowering Diagnosis of Children's Mathematical Cognition through Open-ended Multimodal Tasks
CHI '26· Intelligent Tutoring Systems & Learning Analytics +3
Based on Jaccard similarity of research subtopics & professions (≥60%)