SwapVid: Integrating Video Viewing and Document Exploration with Direct Manipulation

Interactive Data VisualizationData StorytellingContext-Aware ComputingSoftware Engineers & DevelopersUI/UX DesignersHCI Researchers

Document Title

SwapVid: Integrating Video Viewing and Document Exploration with Direct Manipulation

Document Information

  • Domain: Human-Computer Interaction and User Interface Design
  • Keywords: Document and video interaction, video-document matching, lecture videos, screen-shared documents, user interface design

Research Background and Problem Statement

  • Identified Problems or Challenges:

    1. Interfaces combining documents and videos (e.g., slides and videos in online lectures) can enhance content comprehension, but viewers often need to frequently shift focus to navigate between document and video content, increasing cognitive load.
    2. Current tools (e.g., Zoom and Coursera) provide limited support for simultaneous exploration of video and document content, lacking smooth interaction strategies.
  • Importance: The widespread adoption of online learning and communication driven by the COVID-19 pandemic has exacerbated user experience issues related to synchronized document and video content. Addressing these issues can significantly improve users' comprehension efficiency for this type of content.

  • Research Motivation and Related Work:

    1. Existing tools predominantly use side-by-side interfaces, making operations between video and document cumbersome.
    2. Previous studies have attempted direct manipulation for video navigation (e.g., dragging objects in videos), but these approaches were limited to specific scenarios and did not deeply consider seamless transitions between video and document content.
    3. This paper explores how to deeply integrate video and document content to provide a more seamless and intuitive interaction method.

Solution

  • Proposed Method:

    1. Introduce the SwapVid interface—a system that integrates video and document views, enabling seamless switching between video and document modes through direct user operations (e.g., document scrolling, video timeline manipulation).
    2. Utilize OCR (Optical Character Recognition)-based algorithms to automatically establish mapping relationships between video frames and document content.
  • Innovations:

    1. Eliminates the need for split-screen interfaces by enabling single-window interaction through overlay and switching.
    2. Dynamically synchronizes the video timeline with specific content in the document.
    3. Provides a smooth user experience for transitions between document and video modes.
  • Implementation Steps and Key Technologies:

    1. Sequence Analyzer: Uses OCR technology to extract text and its location from video frames and documents, establishing matching relationships between the two.
    2. Interactive Interface Design: Enables mode switching triggered by scrolling and timeline operations, enhancing task flow continuity.
    3. Visualization Features:
      • Highlights time points in the video timeline that match the document content.
      • Highlights the corresponding location in the document view for the currently playing video content.

Research Outcomes

  • Specific Results:

    1. Developed a SwapVid prototype and evaluated it through user experiments with 20 participants.
    2. Experiments demonstrated that SwapVid significantly reduced user time and physical workload in video-based document navigation and document-based video exploration tasks.
    3. Users showed a preference for SwapVid over traditional split-screen interfaces in content exploration tasks.
  • Advantages Compared to Existing Solutions:

    1. Seamless switching reduces the operational cost of split-screen transitions.
    2. OCR enables precise recognition and synchronization of document and video content, improving navigation efficiency.
    3. Experimental results indicate that SwapVid significantly enhances task performance, especially for slide-based documents.
  • Experiment or Evaluation Results:

    1. Navigation Performance: In video-based document exploration tasks (V2D Task), SwapVid significantly reduced task completion time and user mouse movement distance compared to traditional split-screen interfaces.
    2. Subjective Evaluation: Users rated SwapVid's system usability (SUS) as "good" and expressed a preference for it in content exploration tasks.
    3. Cognitive Load: Eye-tracking experiments revealed that cognitive load was relatively lower when using SwapVid.
  • Limitations and Future Directions:

    1. SwapVid currently only supports pre-recorded videos and does not yet accommodate live video (e.g., Zoom meetings).
    2. OCR-based text recognition may face limitations with low-resolution videos or documents without textual content.
    3. User feedback for interface improvements includes dynamic zooming, space-saving picture-in-picture functionality, and better support for mobile devices.
    4. Future work includes integrating the system into live video tools, expanding content matching algorithms to support more document types (e.g., animated slides) and linked documents.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147489/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642515
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Interactive Data Visualization, Data Storytelling, Context-Aware Computing
work
Professions
Software Engineers & Developers, UI/UX Designers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers