The Invisible Mentor: Inferring User Actions from Screen Recordings to Recommend Better Workflows
Authors
Paper Title
The Invisible Mentor: Inferring User Actions from Screen Recordings to Recommend Better Workflows
Publication Info
- Topic area: AI-driven workflow optimization in feature-rich tools like spreadsheets.
- Keywords: Vision-language models, task reflection, workflow optimization, screen recordings, user behavior analysis, AI assistants, Excel, efficiency, generative AI, proactive recommendations.
Background and Problem
- Problem / challenge: Users of complex tools like Excel often fail to discover or utilize efficient features, relying on repetitive, error-prone workflows. Existing solutions require either user prompts or application-specific logging, which are limited in accessibility and usability.
- Significance: Improving workflow efficiency can save time, reduce errors, and enhance user mastery of tools, benefiting both novice and experienced users in professional and personal contexts.
- Motivation and related work: Prior work has explored feature discoverability, action-based suggestions, and AI-driven task assistance. However, these approaches often depend on structured logs, user-initiated prompts, or limited application contexts, leaving a gap in systems that can infer inefficiencies directly from visual evidence.
Solution
- Proposed approach: InvisibleMentor, a system that uses vision-language models (VLMs) to analyze screen recordings, reconstruct user actions, and generate actionable, personalized workflow recommendations via a large language model (LLM).
- Novelty:
- Introduces "vision-grounded task reflection," enabling AI to infer inefficiencies from visual behavior without requiring logs or prompts.
- Develops a two-phase pipeline combining VLMs for action reconstruction and LLMs for generating high-fidelity, context-specific recommendations.
- Demonstrates proactive, post-task guidance that reduces user effort and enhances learning of efficient workflows.
- Evaluates the system's effectiveness through technical benchmarks and user studies, showing significant advantages over prompt-based assistants.
- Procedure and key techniques:
- Phase 1: A VLM processes screen recordings sampled every 5 seconds to extract structured task representations, including user actions and spreadsheet context.
- Phase 2: An LLM analyzes these representations to identify inefficiencies and generate step-by-step recommendations with rationales and benefits.
- Recommendations are displayed post-task in a familiar interface, allowing users to reflect on and adopt better workflows.
Results
- Concrete findings:
- The VLM achieved an F1 score of 0.91 in recognizing 14 common spreadsheet actions.
- InvisibleMentor reduced interaction steps by an average of 32.4% and corrected 18 workflow errors across participants.
- 58 out of 60 generated recommendations were executable and effective.
- Advantage over baselines:
- Participants rated InvisibleMentor's suggestions as significantly more useful, clear, and relevant than those from Excel Copilot.
- InvisibleMentor required less effort to use and was preferred by most participants (p < 0.001).
- Experiments / evaluation:
- A technical benchmark evaluated action recognition accuracy using 25 screen-recorded sessions.
- A user study with 20 participants compared InvisibleMentor to Excel Copilot, assessing suggestion quality, user effort, and workflow improvement.
- Suggestions were evaluated for clarity, relevance, and executability.
- Limitations and future work:
- Limited evaluation on complex workflows (e.g., macros, cross-sheet dependencies).
- Missed subtle formatting actions due to VLM sampling strategy.
- Study conducted in controlled settings; real-world applicability needs further testing.
- Future work includes improving fine-grained action detection, enabling real-time suggestions, and expanding to other domains.
Summary
InvisibleMentor leverages vision-language models to analyze screen recordings and provide actionable, post-task recommendations for improving workflows in tools like Excel. It eliminates the need for user prompts or application-specific logs, making assistance more accessible and proactive. The system demonstrated high accuracy in recognizing user actions and significant advantages over prompt-based assistants in user studies, helping participants discover and adopt more efficient workflows. While currently focused on spreadsheets, the approach shows potential for broader application across domains, with future work aimed at enhancing detection capabilities and real-time support.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 88%
Interview-Informed Generative Agents for Product Discovery: A Validation Study
CHI '26· Human-LLM Collaboration +3
- 88%
Just-In-Time Objectives: A General Approach for Specialized AI Interactions
CHI '26· Human-LLM Collaboration +3
- 75%
DALL: Data Labeling via Data Programming and Active Learning Enhanced by Large Language Models
CHI '26· Human-LLM Collaboration +3
- 75%
"Here, Let Me Help": An Empirical Study of User Interventions in Human–Web Agent Collaboration
CHI '26· Human-LLM Collaboration +3
- 75%
Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM Behavior
CHI '26· Human-LLM Collaboration +3
- 75%
Live in the Loop: Rapid Run-time Feedback for Prompts
CHI '26· Human-LLM Collaboration +3
- 71%
An Analytic Model for Time Efficient Personal Hierarchies
CHI '19· User Research Methods (Interviews, Surveys, Observation) +1
- 71%
Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-Creation
CHI '24· Human-LLM Collaboration +1
- 71%
PDFChatAnnotator: A Human-LLM Collaborative Multi-Modal Data Annotation Tool for PDF-Format Catalogs
IUI '24· Human-LLM Collaboration +1
- 71%
An Exploratory Study on How AI Awareness Impacts Human-AI Design Collaboration
IUI '25· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)