Weak-Annotation of HAR Datasets using Vision Foundation Models
Authors
As wearable-based data annotation remains, to date, a tedious, time-consuming task requiring researchers to dedicate substantial time, benchmark datasets within the field of Human Activity Recognition in lack richness and size compared to datasets available within related fields. Recently, vision foundation models such as CLIP have gained significant attention, helping the vision community advance in finding robust, generalizable feature representations. With the majority of researchers within the wearable community relying on vision modalities to overcome the limited expressiveness of wearable data and accurately label their to-be-released benchmark datasets offline, we propose a novel, clustering-based annotation pipeline to significantly reduce the amount of data that needs to be annotated by a human annotator. We show that using our approach, the annotation of centroid clips suffices to achieve average labelling accuracies close to 90% across three publicly available HAR benchmark datasets. Using the weakly annotated datasets, we further demonstrate that we can match the accuracy scores of fully-supervised deep learning classifiers across all three benchmark datasets. Code as well as supplementary figures and results are publicly downloadable via github.com/mariusbock/weak_har.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
Servo-Gaussian Model to Predict Success Rates in Manual Tracking: Path Steering and Pursuit of 1D Moving Target
UIST '20· Human Pose & Activity Recognition +1
- 100%
Breathing Life Into Biomechanical User Models
UIST '22· Human Pose & Activity Recognition +1
- 80%
The I in Team: Mining Personal Social Interaction Routine with Topic Models from Long-Term Team Data
IUI '18· Human Pose & Activity Recognition +1
- 80%
Demystifying Reward Design in Reinforcement Learning for Upper Extremity Interaction: Practical Guidelines for Biomechanical Simulations in HCI
UIST '25· Human Pose & Activity Recognition +1
- 75%
Computational Rationality as a Theory of Interaction
CHI '22· Computational Methods in HCI
- 67%
Log2Motion: Biomechanical Motion Synthesis from Touch Logs
CHI '26· Hand Gesture Recognition +2
- 67%
ConverSearch: Supporting Experts in Human Behavior Analysis of Conversational Videos with a Multimodal Scene Search Tool (TIIS)
IUI '26· Human Pose & Activity Recognition +2
- 67%
The Potential of Cognitive Circles to Measure Mental Load
UIST '25· Human Pose & Activity Recognition +2
- 60%
Semi-Automated Coding for Qualitative Research: A User-Centered Inquiry and Initial Prototypes
CHI '18· User Research Methods (Interviews, Surveys, Observation) +1
- 60%
Evaluating HCI Research beyond Usability
CHI '18· Mental Health Apps & Online Support Communities +1
Based on Jaccard similarity of research subtopics & professions (≥60%)