Identifying Multimodal Context Awareness Requirements for Supporting User Interaction with Procedural Videos
Authors
Title of the Paper
Identifying Multimodal Context Awareness Requirements for Supporting User Interaction with Procedural Videos
Paper Information
- Subject Area: Human-Computer Interaction, Artificial Intelligence, Context-Aware Technology
- Keywords: Multimodal Context Awareness, Procedural Videos, User Interaction, AI Assistants, Voice User Interface, Video Navigation, Task Alignment
Research Background and Problem
- Identified Problem or Challenge: The study reveals that users frequently switch between the primary context of task execution and the secondary context of interpreting video instructions when watching and imitating tutorial videos. This process can lead to cognitive overload. Furthermore, there has been limited exploration of how to support such continuous context switching, and current intelligent assistant designs fail to effectively consider users' environments or task progress.
- Significance of the Research: With the widespread popularity of procedural "how-to" videos (e.g., cooking and fitness videos), developing intelligent systems that support task alignment can not only enhance user experience but also hold potential for replacing traditional human guidance.
- Motivation and Related Work:
- Existing research primarily focuses on video Q&A functionality and video manipulation technologies, often relying on video metadata and lacking awareness of user tasks and environments.
- For complex tasks such as cooking, users face multiple context switches (e.g., time constraints, environmental state awareness), necessitating the support of context-aware intelligent assistants.
Proposed Solution
- Proposed Method or Solution:
- The authors designed a "Wizard-of-Oz study" to simulate an ideal hands-free interactive system (HFIAI) capable of understanding users' environments and task progress while providing multifaceted support.
- The experiment involved 30 participants performing kitchen cooking tasks, during which all interaction behaviors were recorded through video navigation, reminders, and Q&A functionalities.
- Innovative Aspects:
- By combining user task completion processes with interaction patterns, the study reveals how user characteristics and task contexts influence interaction needs.
- The experiment bypasses current technological limitations (e.g., errors in visual analysis or natural language understanding) to explore the potential capabilities of a truly ideal multimodal system.
- Implementation Steps and Techniques:
- Analyzing the temporal flow of the experiment (from ingredient preparation to serving the dish) to extract data addressing the following questions:
- Types of queries and their timing.
- Motivations behind queries and their association with user backgrounds and personal characteristics.
- How system response capabilities can extend beyond the scope of video metadata.
- Analyzing the temporal flow of the experiment (from ingredient preparation to serving the dish) to extract data addressing the following questions:
Research Outcomes
- Specific Findings:
- User interaction patterns exhibit consistent query timing nodes, but the semantics and motivations of queries vary based on user characteristics.
- Certain interactions require advanced natural language understanding or capabilities beyond video metadata, such as environmental sound awareness, visual capture, and reminder settings.
- Three user personas (process-oriented users, task-oriented users, and multi-tasking users) were identified, with their respective interaction needs and system design implications explored.
- Advantages Compared to Existing Solutions:
- More comprehensive context-aware capabilities that support fine-grained interactions, such as ingredient substitution suggestions and step-specific timing and temperature guidance.
- The design highlights the potential of multimodal assistants while clarifying how learning user query patterns and semantics can enable efficient task support.
- Experiment and Evaluation Results:
- On average, each participant interacted with HFIAI 52 times.
- Consistent query trigger flows were observed among participants (e.g., pausing or playing videos after completing each cooking step).
- Users required not only responses to codified questions (e.g., video navigation commands) but also assistance in understanding cross-domain information and performing environment-aware tasks.
- Limitations and Future Directions:
- The current study focuses on middle-aged and young users, excluding groups unfamiliar with HFIAI or expert chefs.
- The research is centered on cooking tasks; future studies should explore whether similar design principles apply to other task types, such as musical instrument learning or furniture assembly.
- Expanding the system to include multimodal sensing (e.g., smell or taste) and further personalization adjustments could be key areas for future focus.
Conclusion
This paper advances the study of user interaction with procedural videos by progressively demonstrating how multimodal intelligent assistance systems can balance context awareness and task alignment needs. By observing the diversity and commonality in user-system interactions, the authors provide practical guidance for future intelligent assistant design and identify directions for improvements in voice, visual, and environmental sensing technologies.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can user interaction with procedural videos (e.g., tutorial videos) be supported in multimodal context-aware environments?Category: Tutorial Authoring and Skill InstructionSimilar questionsarrow_forward
- How do users' background characteristics and personal behaviors influence their information needs during task execution?Category: Tutorial Authoring and Skill InstructionSimilar questionsarrow_forward
- What technical capabilities should an ideal multimodal system possess to support task alignment (e.g., step matching)?Category: Tutorial Authoring and Skill InstructionSimilar questionsarrow_forward
Practical Problems
1- Users frequently switching between tasks and video contexts while imitating tutorial videos easily experience cognitive overload.Category: Tutorial Authoring and Skill InstructionSimilar questionsarrow_forward
- 67%
Cooking With Agents: Designing Context-aware Voice Interaction
CHI '24· Voice User Interface (VUI) Design +1
- 67%
Supporting Mobile Reading While Walking with Automatic and Customized Font Size Adaptations
CHI '25· Voice User Interface (VUI) Design +1
- 60%
Investigating Context-Aware Collaborative Text Entry on Smartphones using Large Language Models
CHI '25· Voice User Interface (VUI) Design +2
Based on Jaccard similarity of research subtopics & professions (≥60%)