HelpViz: Automatic Generation of Contextual Visual Mobile Tutorials from Text-Based Instructions
Document Title
HelpViz: Automatic Generation of Contextual Visual Mobile Tutorials from Text-Based Instructions
Document Information
- Research Area: Human-Computer Interaction and Automatic Generation of Mobile Application Tutorials
- Keywords: Mobile tutorials, automatic generation, contextual help, deep learning, graphical tutorials, user interaction design, simulated execution, tutorial generation, Android applications, accessibility
Research Background and Problem
- Challenges Identified by Authors: Mobile application tutorials are typically created by professionals, which is costly and time-consuming. Additionally, these tutorials need constant updates to align with software version changes. Many online tutorials exist in text format, requiring users to manually match instructions with device interfaces, which is time-consuming and prone to errors.
- Significance: Automating tutorial generation can reduce costs and quickly adapt tutorials following software updates. This helps mobile users navigate applications more effectively and reduces confusion.
- Motivation and Related Work: Previous studies have explored tutorial creation through expert demonstrations or crowdsourcing, but human involvement is still required. Literature reviews indicate that visual tutorials enhance comprehension and usability, but existing technologies rely on manual creation or semi-automatic generation, failing to achieve full automation.
Solution
- Methods and Approach:
- Text Parsing: Utilize pre-trained deep learning models to parse text instructions and extract actions and target objects.
- Step-by-Step Instruction Segmentation: Decompose instructions into multiple steps based on parsing results.
- Simulated Execution of Instructions: Automatically execute instructions in an array of Android emulators to capture screenshots, video clips, and metadata.
- Tutorial Synthesis: Combine text instructions with generated visual assets (e.g., screenshots and videos) to create interactive graphical tutorials.
- Real-Time Contextual Service: Dynamically display tutorials on mobile devices based on user actions, highlighting the current step being executed.
- Innovations:
- Parse text instructions using deep models to transform online text into graphical tutorials.
- Optimize the parsing model with "beam search" to provide multiple alternatives for potential errors.
- Employ "look-ahead" techniques to address mismatches between instructions and emulator interfaces, maximizing execution success rates.
- Implementation Steps and Key Technologies:
- Pre-trained instruction parsing models for action sequence generation.
- Batch execution of instructions across multiple emulators, handling screen loading issues with dynamic element waiting and scrolling mechanisms.
- Combine screenshots, videos, and metadata during tutorial synthesis to present operation interfaces in a multi-view format.
Research Outcomes
- Specific Results:
- Proposed the first complete pipeline for converting text instructions into visual tutorials automatically.
- Developed a comprehensive system from text parsing to emulator execution and context-sensitive tutorial presentation.
- User testing results showed a preference for HelpViz graphical tutorials, particularly in features like auto-scrolling and step highlighting.
- Advantages:
- Compared to traditional text tutorials, HelpViz-generated visual tutorials are easier to understand, supported by screenshots, and better suited to mobile device user needs.
- Automatically generated tutorials reduce reliance on professional expertise and address challenges posed by software version updates.
- Experimental and Evaluation Results:
- Pipeline Performance Evaluation: The system parsed and batch-converted 187 multi-step instructions from the PixelHelp dataset, achieving a completion rate of 80.9%, significantly higher than the baseline (67.1%).
- User Study: Ten participants compared two tutorial formats, favoring HelpViz, with subjective ratings significantly higher than traditional text tutorials.
- Technical Validation: "Beam search" and "look-ahead" methods significantly improved parsing accuracy and execution success rates.
- Limitations and Future Directions:
- Tutorial generation failures are primarily caused by severe mismatches between instruction text and application interfaces or unsupported emulator environments.
- Does not support complex workflows such as conditional branches or loops, requiring further expansion of parsing model capabilities.
- Optimize tutorials for different devices and versions to address Android fragmentation issues.
- Explore possibilities for adapting desktop software workflows and optimizing tutorial presentation for specific user groups, such as adding voice guidance for visually impaired users.
Through the research on HelpViz, the authors demonstrated the feasibility of large-scale automated generation of visual tutorials, providing higher-quality interactive help resources for mobile users while reducing tutorial maintenance costs. This novel approach offers strong support and inspiration for future automation in human-computer interaction tutorial processes.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can context-relevant visual mobile tutorials be automatically generated from text-based instructions?Category: Paper Reading and Knowledge ExtractionSimilar questionsarrow_forward
- How can deep learning models parse text to extract actions and target objects and generate multi-step tutorials?Category: Paper Reading and Knowledge ExtractionSimilar questionsarrow_forward
- How can mobile device tutorials be dynamically presented to respond to user actions in real time and highlight current steps?Category: Paper Reading and Knowledge ExtractionSimilar questionsarrow_forward
Practical Problems
1- Users spend significant time matching text instructions to mobile device interfaces.Category: Paper Reading and Knowledge ExtractionSimilar questionsarrow_forward
- 100%
Table Illustrator: Puzzle-based interactive authoring of plain tables
CHI '24· Interactive Data Visualization +1
- 100%
Unpacking Visual Metaphors in Infographics: A Design Space
CHI '26· Data Storytelling +1
- 100%
Ragged Blocks: Rendering Structured Text With Style
UIST '25· Interactive Data Visualization +1
- 80%
Exploring Visual Information Flows in Infographics
CHI '20· Interactive Data Visualization +1
- 80%
Pluto: Authoring Semantically Aligned Text and Charts for Data-Driven Communication
IUI '25· Interactive Data Visualization +1
- 75%
Interactive Document Clustering Revisited: A Visual Analytics Approach
IUI '18· Interactive Data Visualization
- 75%
Data-centric disambiguation for data transformation with programming-by-example
IUI '21· Interactive Data Visualization
- 67%
DataToon: Drawing Dynamic Network Comics With Pen + Touch Interaction
CHI '19· Interactive Data Visualization +2
- 67%
SwapVid: Integrating Video Viewing and Document Exploration with Direct Manipulation
CHI '24· Interactive Data Visualization +2
- 60%
What's the Difference?: Evaluating Variations of Multi-Series Bar Charts for Visual Comparison Tasks
CHI '18· Interactive Data Visualization +1
Based on Jaccard similarity of research subtopics & professions (≥60%)