Identifying Multimodal Context Awareness Requirements for Supporting User Interaction with Procedural Videos

Voice User Interface (VUI) DesignContext-Aware ComputingConsumers & Shoppers

Title of the Paper

Identifying Multimodal Context Awareness Requirements for Supporting User Interaction with Procedural Videos

Paper Information

  • Subject Area: Human-Computer Interaction, Artificial Intelligence, Context-Aware Technology
  • Keywords: Multimodal Context Awareness, Procedural Videos, User Interaction, AI Assistants, Voice User Interface, Video Navigation, Task Alignment

Research Background and Problem

  • Identified Problem or Challenge: The study reveals that users frequently switch between the primary context of task execution and the secondary context of interpreting video instructions when watching and imitating tutorial videos. This process can lead to cognitive overload. Furthermore, there has been limited exploration of how to support such continuous context switching, and current intelligent assistant designs fail to effectively consider users' environments or task progress.
  • Significance of the Research: With the widespread popularity of procedural "how-to" videos (e.g., cooking and fitness videos), developing intelligent systems that support task alignment can not only enhance user experience but also hold potential for replacing traditional human guidance.
  • Motivation and Related Work:
    • Existing research primarily focuses on video Q&A functionality and video manipulation technologies, often relying on video metadata and lacking awareness of user tasks and environments.
    • For complex tasks such as cooking, users face multiple context switches (e.g., time constraints, environmental state awareness), necessitating the support of context-aware intelligent assistants.

Proposed Solution

  • Proposed Method or Solution:
    • The authors designed a "Wizard-of-Oz study" to simulate an ideal hands-free interactive system (HFIAI) capable of understanding users' environments and task progress while providing multifaceted support.
    • The experiment involved 30 participants performing kitchen cooking tasks, during which all interaction behaviors were recorded through video navigation, reminders, and Q&A functionalities.
  • Innovative Aspects:
    • By combining user task completion processes with interaction patterns, the study reveals how user characteristics and task contexts influence interaction needs.
    • The experiment bypasses current technological limitations (e.g., errors in visual analysis or natural language understanding) to explore the potential capabilities of a truly ideal multimodal system.
  • Implementation Steps and Techniques:
    • Analyzing the temporal flow of the experiment (from ingredient preparation to serving the dish) to extract data addressing the following questions:
      1. Types of queries and their timing.
      2. Motivations behind queries and their association with user backgrounds and personal characteristics.
      3. How system response capabilities can extend beyond the scope of video metadata.

Research Outcomes

  • Specific Findings:
    • User interaction patterns exhibit consistent query timing nodes, but the semantics and motivations of queries vary based on user characteristics.
    • Certain interactions require advanced natural language understanding or capabilities beyond video metadata, such as environmental sound awareness, visual capture, and reminder settings.
    • Three user personas (process-oriented users, task-oriented users, and multi-tasking users) were identified, with their respective interaction needs and system design implications explored.
  • Advantages Compared to Existing Solutions:
    • More comprehensive context-aware capabilities that support fine-grained interactions, such as ingredient substitution suggestions and step-specific timing and temperature guidance.
    • The design highlights the potential of multimodal assistants while clarifying how learning user query patterns and semantics can enable efficient task support.
  • Experiment and Evaluation Results:
    • On average, each participant interacted with HFIAI 52 times.
    • Consistent query trigger flows were observed among participants (e.g., pausing or playing videos after completing each cooking step).
    • Users required not only responses to codified questions (e.g., video navigation commands) but also assistance in understanding cross-domain information and performing environment-aware tasks.
  • Limitations and Future Directions:
    • The current study focuses on middle-aged and young users, excluding groups unfamiliar with HFIAI or expert chefs.
    • The research is centered on cooking tasks; future studies should explore whether similar design principles apply to other task types, such as musical instrument learning or furniture assembly.
    • Expanding the system to include multimodal sensing (e.g., smell or taste) and further personalization adjustments could be key areas for future focus.

Conclusion

This paper advances the study of user interaction with procedural videos by progressively demonstrating how multimodal intelligent assistance systems can balance context awareness and task alignment needs. By observing the diversity and commonality in user-system interactions, the authors provide practical guidance for future intelligent assistant design and identify directions for improvements in voice, visual, and environmental sensing technologies.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/96369/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581006
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Voice User Interface (VUI) Design, Context-Aware Computing
work
Professions
Consumers & Shoppers
article
Content Status
Full text indexed
hub
Related Papers
3 related papers