ReVisor: A Reflective Design Tool for Instructional Designers to Improve Teacher Training Materials via AI Discussions

Human-LLM CollaborationAI-Assisted Decision-Making & AutomationIntelligent Tutoring Systems & Learning AnalyticsOnline Course DesignersUniversity Professors & Researchers

Paper Title

ReVisor: A Reflective Design Tool for Instructional Designers to Improve Teacher Training Materials via AI Discussions

Publication Info

  • Topic area: AI-driven tools for improving instructional design using classroom discourse analysis.
  • Keywords: instructional design, classroom discourse, teacher training, AI feedback, multi-agent systems, ambiguity detection, reflective practice, LLMs, educational technology, iterative refinement.

Background and Problem

  • Problem / challenge: Instructional designers lack systematic tools to evaluate and refine teacher training materials based on real classroom discourse, which is qualitative, voluminous, and difficult to analyze. Existing tools primarily summarize data without offering actionable insights for iterative refinement.
  • Significance: Improving teacher training materials is critical for bridging the gap between pedagogical theory and classroom practice, ensuring that teachers can effectively implement instructional strategies.
  • Motivation and related work: Prior work includes manual annotation, visualization tools, multimodal sensing, and AI-based feedback systems. However, these approaches focus on summarizing data or assisting teachers, not on supporting instructional designers in refining materials based on real-world enactment data.

Solution

  • Proposed approach: ReVisor, a reflective design tool that uses multi-agent discussions powered by large language models (LLMs) to analyze classroom transcripts, identify ambiguous applications of teaching strategies, and provide actionable revision suggestions.
  • Novelty:
    1. Multi-agent reasoning pipeline to detect ambiguity in classroom discourse and classify teaching method applications as clear, ambiguous, or absent.
    2. Integration of AI-driven feedback loops into iterative instructional design workflows.
    3. Development of a user interface that combines statistical summaries, AI-agent discussions, and improvement suggestions to guide instructional designers.
    4. Release of an open-source dataset of AI-agent discussions for further research on multi-agent reasoning in instructional design.
  • Procedure and key techniques:
    1. Analyze classroom transcripts using two LLM agents that independently label teacher utterances as Present, Absent, or Uncertain based on training materials.
    2. Trigger agent discussions when disagreement occurs, with up to three conversational turns to resolve ambiguity.
    3. Classify utterances as Clear, Ambiguous, or Absent based on agent consensus or unresolved disagreement.
    4. Extract improvement suggestions for ambiguous cases to refine training materials.
    5. Provide iterative feedback through a user interface, allowing designers to revise materials and reanalyze their impact.

Results

  • Concrete findings:
    • The pipeline achieved an F1 score of 0.73 for detecting teaching method presence.
    • Ambiguous utterances identified by AI were rated significantly lower by human participants, validating AI disagreement as a proxy for pedagogical ambiguity.
    • ReVisor reduced the number of ambiguous cases in training materials by an average of 10.89 per revision cycle.
  • Advantage over baselines:
    • Participants using ReVisor produced more specific, evidence-based, and actionable revisions compared to a baseline system.
    • ReVisor facilitated the integration of theoretical methods with practical examples, resulting in clearer and better-structured training materials.
  • Experiments / evaluation:
    • A user study with 10 professional instructional designers showed that ReVisor significantly improved the ease of refining materials (p = 0.039) and supported the inclusion of practical examples (p = 0.031).
    • A validation study with 34 undergraduate students confirmed that AI disagreement aligns with human judgments of pedagogical ambiguity (t-values ranging from -2.19 to -9.14, all p < 0.05).
  • Limitations and future work:
    • Small sample size for the user study and limited scope to talk moves within teacher training materials.
    • Dependence on static classroom transcripts and mid-scale LLMs, which may limit generalizability.
    • Future work includes exploring heterogeneous LLM configurations, expanding to other pedagogical domains, and integrating ReVisor into teacher training workshops.

Summary

ReVisor is an AI-driven tool that helps instructional designers refine teacher training materials by analyzing classroom discourse and identifying areas of ambiguity through multi-agent LLM discussions. It provides actionable feedback, supports iterative refinement, and bridges the gap between theory and practice. The system was validated through technical evaluations and user studies, demonstrating its ability to improve the clarity and applicability of training materials. While limitations exist, ReVisor offers a promising framework for integrating AI into instructional design workflows, with potential applications across diverse educational contexts.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222573/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790456
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation, Intelligent Tutoring Systems & Learning Analytics
work
Professions
Online Course Designers, University Professors & Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers