TutoAI: a cross-domain framework for AI-assisted mixed-media tutorial creation on physical tasks

AI-Assisted Creative WritingCreative Collaboration & Feedback SystemsComputational Methods in HCISoftware Engineers & DevelopersUI/UX DesignersFreelancers (Design, Writing, Translation)

Document Title

TutoAI: A Cross-domain Framework for AI-assisted Mixed-media Tutorial Creation on Physical Tasks

Document Information

  • Subject Area: Human-Computer Interaction (HCI), cross-domain mixed-media tutorial creation, AI-assisted content generation
  • Keywords: Human-Computer Interaction, mixed-media tutorials, AI-assisted creation, large language models, video processing techniques, user interface design, step and object extraction, dependencies, cross-domain generalization

Research Background and Problem

  • Problems and Challenges:

    • Video tutorials, as an important source for learning skills, often rely on linear timeline navigation, making it difficult to browse and locate specific information, leading users to miss critical operational details.
    • Creating high-quality mixed-media tutorials (integrating video, images, text, and diagrams) is a key approach to achieving better navigability and usability, but manual creation of such tutorials is time-consuming and labor-intensive.
    • Existing automated methods are often designed for specific domains, making it difficult to generalize across domains and lacking a shared methodological framework for cross-domain applications.
    • AI (especially large language models, LLMs) offers new possibilities, but effectively applying it to multimodal data processing (e.g., combining video and text) remains an unresolved challenge.
  • Significance:

    • The navigability and multidimensional information representation of mixed-media tutorials provide users with a more efficient learning experience, addressing navigation issues in traditional video tutorials.
    • Developing a cross-domain framework can overcome domain-specific limitations and provide universal tools and methods for creating mixed-media tutorials across different fields.
  • Research Motivation and Related Work:

    • Previous studies have proposed methods for extracting mixed-media tutorials, but most focus on specific domains (e.g., makeup, cooking, and software tutorials).
    • Some studies have attempted to explore cross-domain methods but lack a systematic framework to integrate and evaluate multimodal AI models and user interaction techniques.
    • The authors argue for the need for a comprehensive cross-domain framework, including common tutorial components, applicable AI models and evaluation methods, as well as user-facing interfaces for content review.

Solution

  • Proposed Framework:

    • The authors propose TutoAI, a framework that achieves AI-assisted cross-domain mixed-media tutorial creation through three levels:
      1. Component Level: Identifies key components of mixed-media tutorials (steps, objects, and dependencies) and defines their cross-domain representations.
      2. Model Level: Proposes a systematic method for selecting, assembling, and evaluating AI models to extract tutorial components.
      3. User Interface Level: Introduces design guidelines for user interfaces, enabling users to review and modify AI-generated content.
  • Innovations:

    • Introduced a cross-domain classification framework for mixed-media tutorial components.
    • Designed a systematic process for assembling AI models, integrating the advantages of multimodal data (e.g., text, video, images).
    • Reduced users' cognitive load through user interface design, allowing efficient correction of AI-generated results.
  • Implementation Steps and Key Technologies:

    • Step Extraction: Combined large language models (GPT-3.5 for text transcription and summarization) with video segmentation detectors (for generating timestamps and thumbnails).
    • Object Extraction: Used large language models to extract object names from text and open-vocabulary object detectors (e.g., OWL-ViT) to annotate objects in video frames.
    • Dependency Construction: Built dependency graphs by analyzing shared objects between steps and generated visualizations of these dependencies.
    • User Interface Design: Developed a component-review-based step-by-step interactive workflow, enabling users to iteratively refine the generated tutorial components.

Research Outcomes

  • Specific Results:

    • Proposed a complete cross-domain tutorial creation framework (TutoAI), defined a classification of mixed-media tutorial components, and created a prototype user interface.
    • Built a final AI pipeline for step and object extraction, integrating the capabilities of multiple models (LLMs, object detectors, etc.).
    • Experimental evaluations demonstrated that TutoAI-generated tutorial components significantly outperformed baseline methods (e.g., YouTube's auto-generated video chapters).
  • Advantages over Existing Solutions:

    • Strong cross-domain generalization: The step extraction pipeline and object extraction techniques are not limited to specific domains.
    • Multimodal data support: Simultaneously processes text transcription, video, and images, producing more comprehensive tutorials.
    • User interaction support: Introduced a human-in-the-loop interface, allowing manual optimization of AI-generated results.
  • Experiments or Evaluation Results:

    1. Pipeline Performance:
      • Step boundary detection achieved an average F1 score of 0.59 across datasets from four domains; GPT-based text summarization outperformed traditional models; object extraction achieved an average F1 score of 0.88.
    2. User Study:
      • Among 24 general users, evaluations indicated that TutoAI provided higher-quality tutorial components (e.g., text and thumbnails) compared to YouTube Chapters.
      • Two recruited YouTube creators noted that TutoAI's assistance tools significantly benefited their existing workflows, especially praising the clarity of the dependency graphs.
  • Limitations and Future Directions:

    • Limitations:

      • The current method is primarily suited for physical tasks involving multiple objects and steps (e.g., cooking and hardware assembly) and is less applicable to non-physical tasks (e.g., abstract knowledge explanations).
      • The accuracy of AI-generated results is limited (e.g., dependency parsing may miss certain relationships).
      • The scale of current experiments is small, with testing conducted on only a few domains and user groups.
    • Future Directions:

      • Expand applicability: Optimize dependency modeling and explore a broader range of task types (e.g., purely educational videos).
      • Conduct deeper user experience research: Evaluate the actual learning improvements provided by mixed-media tutorials.
      • Improve model accuracy: Integrate multimodal LLM technologies and multitask models to enhance component extraction capabilities.
      • Introduce more flexible output design styles to adapt generated tutorials to diverse scenarios and user needs.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147303/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642443
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
AI-Assisted Creative Writing, Creative Collaboration & Feedback Systems, Computational Methods in HCI
work
Professions
Software Engineers & Developers, UI/UX Designers, Freelancers (Design, Writing, Translation)
article
Content Status
Full text indexed
hub
Related Papers
3 related papers