MonoEye: Multimodal Human Motion Capture System Using A Single Ultra-Wide Fisheye Camera
Authors
We present MonoEye, a multimodal human motion capture system using a single RGB camera with an ultra-wide fisheye lens, mounted on the user’s chest. Existing optical motion capture systems use multiple cameras, which are synchronized and require camera calibration. These systems also have usability constraints that limit the user’s movement and operating space. Since the MonoEye system is based on a wearable single RGB camera, the wearer’s 3D body pose can be captured without space and environment limitations. The body pose, captured with our system, is aware of the camera orientation and therefore it is possible to recognize various motions that existing egocentric motion capture systems cannot recognize. Furthermore, the proposed system captures not only the wearer’s body motion but also their viewport using the head pose estimation and an ultra-wide image. To implement robust multimodal motion capture, we design three deep neural networks: BodyPoseNet, HeadPoseNet, and CameraPoseNet, that estimate 3D body pose, head pose, and camera pose in real-time, respectively. We train these networks with our new extensive synthetic dataset providing 680K frames of renderings of people with a wide range of body shapes, clothing, actions, backgrounds, and lighting conditions. To demonstrate the interactive potential of the MonoEye system, we present several application examples from common body gestural to context-aware interactions.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
OddEyeCam: A Sensing Technique for Body-Centric Peephole Interaction using WFoV RGB and NFoV Depth Cameras
UIST '20· Full-Body Interaction & Embodied Input +2
- 67%
How Relevant are Incidental Power Poses for HCI?
CHI '18· Full-Body Interaction & Embodied Input +1
- 67%
KeyTime: Super-Accurate Prediction of Stroke Gesture Production Times
CHI '18· Full-Body Interaction & Embodied Input +1
- 67%
The Dissimilarity-Consensus Approach to Agreement Analysis in Gesture Elicitation Studies
CHI '19· Full-Body Interaction & Embodied Input +1
- 67%
Predicting Mid-Air Interaction Movements and Fatigue Using Deep Reinforcement Learning
CHI '20· Full-Body Interaction & Embodied Input +1
- 67%
Pose-on-the-Go: Approximating User Pose with Smartphone Sensor Fusion and Inverse Kinematics
CHI '21· Full-Body Interaction & Embodied Input +1
- 67%
Towards designing for everyday embodied remembering: Findings from a diary study
DIS '23· Full-Body Interaction & Embodied Input +1
- 67%
Below the Surface: Unobtrusive Activity Recognition for Work Surfaces using RF-radar sensing
IUI '18· Human Pose & Activity Recognition +1
- 67%
Assisting Group Activity Analysis through Hand Detection and Identification in Multiple Egocentric Videos
IUI '19· Full-Body Interaction & Embodied Input +1
- 67%
ConvBoost: Boosting ConvNets for Sensor-based Activity Recognition
UbiComp '23· Human Pose & Activity Recognition +1
Based on Jaccard similarity of research subtopics & professions (≥60%)