MMTSA: Multi-Modal Temporal Segment Attention Network for Efficient Human Activity Recognition
Authors
"Multimodal sensors provide complementary information to develop accurate machine-learning methods for human activity recognition (HAR), but introduce significantly higher computational load, which reduces efficiency. This paper proposes an efficient multimodal neural architecture for HAR using an RGB camera and inertial measurement units (IMUs) called Multimodal Temporal Segment Attention Network (MMTSA). MMTSA first transforms IMU sensor data into a temporal and structure-preserving gray-scale image using the Gramian Angular Field (GAF), representing the inherent properties of human activities. MMTSA then applies a multimodal sparse sampling method to reduce data redundancy. Lastly, MMTSA adopts an inter-segment attention module for efficient multimodal fusion. Using three well-established public datasets, we evaluated MMTSA's effectiveness and efficiency in HAR. Results show that our method achieves superior performance improvements (11.13% of cross-subject F1-score on the MMAct dataset) than the previous state-of-the-art (SOTA) methods. The ablation study and analysis suggest that MMTSA's effectiveness in fusing multimodal data for accurate HAR. The efficiency evaluation on an edge device showed that MMTSA achieved significantly better accuracy, lower computational load, and lower inference latency than SOTA methods." https://doi.org/10.1145/3610872
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
Rediscovering Affordance: A Reinforcement Learning Perspective
CHI '22· Human Pose & Activity Recognition
- 100%
How much Unlabeled Data is Really Needed for Effective Self-Supervised Human Activity Recognition?
UbiComp '23· Human Pose & Activity Recognition
- 100%
On the Utility of Virtual On-body Acceleration Data for Fine-grained Human Activity Recognition
UbiComp '23· Human Pose & Activity Recognition
- 100%
MI-Poser: Human Body Pose Tracking Using Magnetic and Inertial Sensor Fusion with Metal Interference Mitigation
UbiComp '23· Human Pose & Activity Recognition
- 100%
SF-Adapter: Computational-Efficient Source-Free Domain Adaptation for Human Activity Recognition
UbiComp '24· Human Pose & Activity Recognition
- 100%
PmTrack: Enabling Personalized mmWave-based Human Tracking
UbiComp '24· Human Pose & Activity Recognition
- 100%
XRF55: A Radio Frequency Dataset for Human Indoor Action Analysis
UbiComp '24· Human Pose & Activity Recognition
- 100%
Semantic Loss: A New Neuro-Symbolic Approach for Context-Aware Human Activity Recognition
UbiComp '24· Human Pose & Activity Recognition
- 100%
TS2ACT: Few-Shot Human Activity Sensing with Cross-Modal Co-Learning
UbiComp '24· Human Pose & Activity Recognition
- 100%
IMUGPT 2.0: Language-Based Cross Modality Transfer for Sensor-Based Human Activity Recognition
UbiComp '24· Human Pose & Activity Recognition
Based on Jaccard similarity of research subtopics & professions (≥60%)