AMIR: Active Multimodal Interaction Recognition from Video and Network Traffic in Connected Environments
Authors
Activity recognition using video data is widely adopted for elder care, monitoring for safety and security, and home automation. Unfortunately, using video data as the basis for activity recognition can be brittle, since models trained on video are often not robust to certain environmental changes, such as camera angle and lighting changes. There has been a proliferation of network-connected devices in home environments. Interactions with these smart devices are associated with network activity, making network data a potential source for recognizing these device interactions. This paper advocates for the synthesis of video and network data for robust interaction recognition in connected environments. We consider machine learning-based approaches for activity recognition, where each labeled activity is associated with both a video capture and an accompanying network traffic trace. We develop a simple but effective framework AMIR (Active Multimodal Interaction Recognition)1 that trains independent models for video and network activity recognition respectively, and subsequently combines the predictions from these models using a meta-learning framework. Whether in lab or at home, this approach reduces the amount of "paired" demonstrations needed to perform accurate activity recognition, where both network and video data are collected simultaneously. Specifically, the method we have developed requires up to 70.83% fewer samples to achieve 85% F1 score than random data collection, and improves accuracy by 17.76% given the same number of samples. https://dl.acm.org/doi/10.1145/3580818
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- Can multimodal integration of video and network traffic improve robustness and accuracy of activity recognition in home environments?Category: Contextual Example Selection and Rare Pattern CoverageSimilar questionsarrow_forward
- How can active learning algorithms reduce the cost of obtaining jointly labeled samples?Category: Contextual Example Selection and Rare Pattern CoverageSimilar questionsarrow_forward
- What is the potential of network traffic as an activity recognition data source, and how can it complement video data to improve overall performance?Category: Contextual Example Selection and Rare Pattern CoverageSimilar questionsarrow_forward
Practical Problems
1- Traditional video-based activity recognition is sensitive to environmental changes and has limited accuracy.Category: Contextual Example Selection and Rare Pattern CoverageSimilar questionsarrow_forward
- 67%
Self-Supervised Representation Learning and Temporal-Spectral Feature Fusion for Bed Occupancy Detection
UbiComp '24· Human Pose & Activity Recognition +2
- 60%
Wall++: Room-Scale Interactive and Context-Aware Sensing
CHI '18· Human Pose & Activity Recognition +2
- 60%
Learning About Social Context From Smartphone Data: Generalization Across Countries and Daily Life Moments
CHI '24· Human Pose & Activity Recognition +2
Based on Jaccard similarity of research subtopics & professions (≥60%)