'Show It, Don't Just Say It': The Complementary Effects of Instruction Multimodality for Software Guidance
Authors
Paper Title
"Show It, Don't Just Say It": The Complementary Effects of Instruction Multimodality for Software Guidance
Publication Info
- Topic area: Multimodal instruction strategies for software learning and their implications for AI tutoring systems.
- Keywords: multimodal learning, software guidance, adaptive tutoring, cognitive load, student agency, digital territoriality, spatial precision, temporal precision, Figma, AI tutoring systems.
Background and Problem
- Problem / challenge: Existing software guidance systems often rely on single or limited modalities (e.g., text or video), which fail to address the spatial and temporal precision required for complex software learning. They also lack adaptiveness to student needs and may intrude on learner agency.
- Significance: Effective multimodal guidance can enhance software learning by improving comprehension, reducing cognitive load, and fostering student engagement. Understanding how human teachers balance precision and agency can inform the design of AI tutoring systems.
- Motivation and related work: Prior research highlights the benefits of multimodal learning and adaptive tutoring but lacks insights into how human teachers coordinate modalities (speech, annotations, remote control) in software instruction. This study addresses the gap by analyzing natural teaching behaviors to inform AI system design.
Solution
- Proposed approach: Observational study of human teacher-student pairs to analyze how speech, visual annotations, and remote screen control are used and coordinated during software instruction.
- Novelty:
- Empirical evidence of modality-specific functions and trade-offs in software learning.
- Introduction of precision-agency trade-off and digital territoriality as key design constraints for AI tutoring systems.
- Identification of adaptive teaching strategies for balancing instructional precision and learner autonomy.
- Design recommendations for AI systems, including ghost cursors, fading annotations, and timeline scrubbing.
- Procedure and key techniques:
- Observed 10 teacher-student pairs teaching two Figma tasks (Weather Icons and Weather Cards) using speech, annotations, and remote control.
- Collected and analyzed multimodal data (speech transcripts, visual annotations, and screen control actions).
- Developed coding frameworks to categorize modality usage and assess teaching adaptiveness.
- Conducted semi-structured interviews to capture teacher and student perspectives.
Results
- Concrete findings:
- Speech was foundational for instruction, providing context, commands, and feedback, but lacked spatial and temporal precision.
- Visual annotations enhanced spatial precision, aiding attention and conceptual understanding.
- Remote screen control provided both spatial and temporal precision but intruded on student agency.
- Teachers adapted modality use based on task demands, student responses, and individual teaching philosophies.
- Advantage over baselines:
- Multimodal coordination outperformed single-modality approaches by addressing both spatial and temporal precision while balancing cognitive load and agency.
- Adaptive teaching strategies improved learning outcomes, with all students completing Test 1 and 90% completing Test 2 successfully.
- Experiments / evaluation:
- Conducted in a controlled lab environment with 10 teacher-student pairs.
- Tasks involved creating weather-themed components in Figma, progressing from basic vector manipulation to advanced component design.
- Data included 12.4 hours of recordings, segmented into lesson steps, and analyzed for modality usage patterns.
- Limitations and future work:
- Focused on Figma; findings may not generalize to other domains like programming.
- Controlled lab setting may differ from real-world contexts.
- Sample size (N=10) limits generalizability; future work should explore larger and more diverse populations.
- Future studies should link modality coordination to measurable learning outcomes and test additional modalities (e.g., gaze highlighting).
Summary
This study analyzed how human teachers use and coordinate speech, visual annotations, and remote screen control to teach software skills, revealing modality-specific strengths and trade-offs. Speech provided foundational guidance, annotations enhanced spatial precision, and screen control supported complex demonstrations but intruded on student agency. Teachers adapted their strategies to balance precision and learner autonomy, introducing the concepts of precision-agency trade-off and digital territoriality. These findings inform the design of AI tutoring systems, suggesting features like ghost cursors and fading annotations to improve clarity while preserving student control. Future work should test these insights across broader domains and learning contexts.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
Expanding a Temporal Vocabulary Towards Designing for “Undiscoverable Learning”
CHI '26· Intelligent Tutoring Systems & Learning Analytics +2
- 83%
EvaluAId: Human-AI Collaborative Evaluation of Open-Ended Student Essays
CHI '26· Human-LLM Collaboration +3
- 71%
"Listen to the Teachers": Research-Based Personas for Translating Classroom Realities into Actionable HCI Design
CHI '26· User Research Methods (Interviews, Surveys, Observation) +3
- 71%
OpenCD: Empowering Diagnosis of Children's Mathematical Cognition through Open-ended Multimodal Tasks
CHI '26· Intelligent Tutoring Systems & Learning Analytics +3
- 67%
Toward Automated Feedback on Teacher Discourse to Enhance Teacher Learning
CHI '20· Intelligent Tutoring Systems & Learning Analytics +1
- 67%
Exploring the Design and Impact of Interactive Worked Examples for Learners with Varying Prior Knowledge
CHI '26· Intelligent Tutoring Systems & Learning Analytics +2
Based on Jaccard similarity of research subtopics & professions (≥60%)