RubySlippers: Supporting Content-based Voice Navigation for How-to Videos

Voice User Interface (VUI) DesignIntelligent Voice Assistants (Alexa, Siri, etc.)Conversational ChatbotsUI/UX DesignersHCI Researchers

Document Title

RubySlippers: Supporting Content-based Voice Navigation for How-to Videos

Document Information

  • Research Domain: Human-Computer Interaction, Voice User Interface Design, Video Navigation
  • Keywords: Voice User Interface, Video Navigation, Tutorial Videos, Content Navigation, How-to Videos

Research Background and Problem Statement

  • Identified Problems or Challenges:

    1. When watching how-to videos, users often navigate by dragging the timeline, but for videos involving physical activities, this operation forces frequent switching between video control and actual tasks, increasing cognitive load.
    2. Converting timeline operations directly into voice commands (e.g., "Jump to 3 minutes and 15 seconds") limits users' ability to preview content quickly.
    3. Current challenges in voice navigation include difficulty in remembering available commands, speech recognition errors, and navigation methods inconsistent with users' habitual operations.
  • Significance: Improving voice navigation for how-to videos can help users control video content more efficiently while reducing interruptions caused by video control, offering significant practical value.

  • Research Motivation and Related Work: Existing research on video navigation primarily focuses on improving timeline interfaces or adding thumbnails for content display. However, few studies explore optimizing voice user interfaces to support content-based navigation strategies. Additionally, other fields, such as educational videos, have proposed data-driven interaction techniques, but these are not fully applicable to voice navigation design in how-to videos.

Solution

  • Main Method or Solution: The RubySlippers system proposes keyword-based content navigation through voice commands. Its core is a computational pipeline that automatically detects referable elements in videos and optimizes video segmentation to reduce the number of navigation commands.

  • Innovations:

    1. Introduced keyword-driven content navigation, eliminating the need for users to remember specific timestamps.
    2. Designed a user-friendly navigation interface, including automatic keyword recommendations and query update functionality, addressing users' difficulty in recalling precise terms.
    3. Incorporated an automatic bookmarking feature that marks frequently accessed scenes based on user interactions.
  • Implementation Steps and Key Technologies:

    1. Computational Pipeline Design: Preprocess video subtitles using natural language processing techniques to capture keywords, segment videos based on sentence themes and lexical distribution, and condense segments into referable units.
    2. Navigation Interface Optimization:
      • Keyword Search: Quickly locate scenes by matching keywords.
      • Query Updates: Users can add or remove keywords to refine search results.
      • Recommendation Mechanism: Dynamically generate navigation suggestions based on video content and user behavior.
    3. System Design: Integrate voice recognition, text processing, video interface, and support real-time interaction.

Research Outcomes

  • Specific Results:

    1. Provided a keyword-driven content navigation method, significantly reducing the number of navigation commands used.
    2. Developed the RubySlippers prototype system and evaluated it with 12 participants using metrics such as task completion time, cognitive load, and user satisfaction.
  • Advantages:

    1. Compared to simple timeline voice navigation, users achieved higher navigation efficiency and reduced interaction frequency in multi-goal task scenarios.
    2. User feedback indicated that keyword navigation better meets practical needs, reducing cognitive load and operational pressure.
    3. The automatic bookmarking feature alleviates the burden of repeated navigation.
  • Experiment or Evaluation Results:

    1. Experiments showed a significant reduction in interaction frequency (average decrease of approximately 43%) during multi-goal search tasks across three types of interaction tasks.
    2. NASA-TLX cognitive load assessments revealed noticeable reductions in time demand and frustration when combining content navigation.
    3. Keyword navigation helped users better understand video content and enhanced learning outcomes.
  • Limitations and Future Directions:

    1. Speech Recognition Errors: Current speech recognition technology still limits the navigation experience; future improvements in recognition accuracy are needed.
    2. Dependence on Visual Display: The RubySlippers interface relies on video display, restricting its potential for fully voice-based assistants. Future work should explore screen-independent design solutions.
    3. Domain Generalizability: Videos with fewer keywords and actions (e.g., origami tutorials) limit the system's effectiveness. Future efforts should integrate computer vision techniques to expand applicability to more types of how-to videos.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/47702/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3411764.3445131
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Voice User Interface (VUI) Design, Intelligent Voice Assistants (Alexa, Siri, etc.), Conversational Chatbots
work
Professions
UI/UX Designers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
5 related papers