Pose-on-the-Go: Approximating User Pose with Smartphone Sensor Fusion and Inverse Kinematics
Authors
Document Title
Pose-on-the-Go: Approximating User Pose with Smartphone Sensor Fusion and Inverse Kinematics
Document Information
- Subject Area: Smartphone-based user pose estimation and motion capture
- Keywords: body pose, pose estimation, smartphone, sensor fusion, inverse kinematics, mobile applications, motion capture, human interaction, augmented reality, virtual reality
Research Background and Problem
-
Identified Problems or Challenges:
- Most current full-body motion capture technologies require additional equipment, such as wearable sensors or multi-camera systems, which are costly and limit usage in mobile scenarios.
- Even external devices like Xbox Kinect cannot be conveniently used in non-specific environments (e.g., outdoors).
- There is a lack of self-contained smartphone-based solutions for full-body pose estimation.
-
Significance:
- Achieving full-body pose estimation on standard smartphones could expand interactive experiences in areas such as gaming, social applications, and health monitoring.
- The ubiquity of smartphones and their multi-sensor capabilities make them a highly feasible solution.
-
Research Motivation and Related Work:
- Researchers reviewed previous external sensor and wearable tracking technologies, including optical cameras, IMUs, magnetic fields, and ultrasonic methods.
- The Pose-on-the-Go system was proposed to leverage existing smartphone hardware (e.g., cameras, IMUs, and depth sensors) to enable a self-contained and low-cost mobile motion capture solution.
Solution
-
Proposed Solution:
- The Pose-on-the-Go system achieves smartphone-based full-body pose estimation through multi-sensor fusion and inverse kinematics (IK) techniques.
- It integrates data from the smartphone's front and rear RGB cameras, depth camera, IMU, and touchscreen to generate a real-time animated skeleton of the user.
-
Innovations:
- The first system to achieve full-body pose estimation using a single smartphone without requiring additional hardware or modifications.
- Provides dynamic pose estimation for the head, torso, arms, and legs, which, while approximate, is suitable for applications with lower interaction demands.
- Offers a lightweight software implementation that can theoretically support existing smartphones through software updates.
-
Implementation Steps and Key Techniques:
- Head Position and Orientation: Head tracking is based on the smartphone's front camera and ARKit API.
- Torso Orientation Estimation: The depth camera captures the chest region below the head to calculate torso orientation.
- Arm and Smartphone Motion Estimation: IMU and inverse kinematics are used to generate plausible arm poses.
- Leg and Walking Animation: Combines the smartphone's absolute position in the environment (6-DOF) with predicted motion patterns (e.g., walking or running) to simulate leg movements through IK animation.
- A data synchronization mechanism integrates asynchronous data streams from multiple sensors to optimize real-time performance.
Research Outcomes
-
Specific Outcomes:
- Pose-on-the-Go can capture full-body poses and estimate joint positions with a 3D spatial error of less than 25 cm.
- Experiments demonstrated the system's adaptability to dynamic movements such as raising hands, turning, and walking.
-
Advantages Compared to Existing Solutions:
- Compared to external tracking systems, Pose-on-the-Go is more portable and cost-effective, requiring no additional hardware.
- Compared to similar head-mounted systems, it maximizes accessibility and usability through smartphone sensors.
-
Experimental or Evaluation Results:
- The system achieved an average angular error of 6-10 degrees for head orientation and a positional error of 9-27 cm for major joints (e.g., shoulders, elbows).
- Absolute spatial positioning error was approximately 11 cm.
- User testing indicated that while the system's gait and body pose simulation are approximate, they are sufficient for low-interaction scenarios such as gaming and exercise.
-
Limitations and Future Directions:
- Key limitations include the inability to accurately track unobserved body parts (e.g., the arm not holding the phone or legs) and reliance on estimation for these areas.
- The current implementation has latency (~350 ms), which may impact user experience in some real-time interaction scenarios.
- High power consumption limits usage to approximately 2 hours on an iPhone XR, requiring further optimization.
- Future research could expand on:
- Enhancing tracking of the non-phone-holding arm using smartwatches or additional sensors.
- Improving the precision of touchscreen inputs, such as recognizing finger types and poses.
- Incorporating facial expression and voice dynamics capture to enhance social interaction capabilities.
Conclusion
Pose-on-the-Go introduces an innovative method for full-body pose estimation, leveraging existing smartphone hardware to enable motion capture without additional equipment. This approach reduces the barrier to entry for users and provides a lightweight solution for various applications, including mobile gaming, health tracking, and remote social interaction. However, further improvements in accuracy and real-time performance are needed.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can smartphone multi-sensor fusion and inverse kinematics achieve full-body pose estimation?Category: Human Pose and Skeleton SensingSimilar questionsarrow_forward
- Can smartphones dynamically estimate head, torso, arm, and leg poses using only existing hardware?Category: Human Pose and Skeleton SensingSimilar questionsarrow_forward
- How does the Pose-on-the-Go system provide real-time motion capture while maintaining low cost and portability?Category: Human Pose and Skeleton SensingSimilar questionsarrow_forward
Practical Problems
1- Existing full-body motion capture requires expensive equipment and has limited mobility.Category: Human Pose and Skeleton SensingSimilar questionsarrow_forward
- 100%
How Relevant are Incidental Power Poses for HCI?
CHI '18· Full-Body Interaction & Embodied Input +1
- 100%
KeyTime: Super-Accurate Prediction of Stroke Gesture Production Times
CHI '18· Full-Body Interaction & Embodied Input +1
- 100%
The Dissimilarity-Consensus Approach to Agreement Analysis in Gesture Elicitation Studies
CHI '19· Full-Body Interaction & Embodied Input +1
- 100%
Predicting Mid-Air Interaction Movements and Fatigue Using Deep Reinforcement Learning
CHI '20· Full-Body Interaction & Embodied Input +1
- 100%
Towards designing for everyday embodied remembering: Findings from a diary study
DIS '23· Full-Body Interaction & Embodied Input +1
- 100%
Assisting Group Activity Analysis through Hand Detection and Identification in Multiple Egocentric Videos
IUI '19· Full-Body Interaction & Embodied Input +1
- 67%
Real-time 3D Target Inference via Biomechanical Simulation
CHI '24· Full-Body Interaction & Embodied Input +2
- 67%
MeCap: Whole-Body Digitization for Low-Cost VR/AR Headsets
UIST '19· Full-Body Interaction & Embodied Input +2
- 67%
OddEyeCam: A Sensing Technique for Body-Centric Peephole Interaction using WFoV RGB and NFoV Depth Cameras
UIST '20· Full-Body Interaction & Embodied Input +2
- 67%
MonoEye: Multimodal Human Motion Capture System Using A Single Ultra-Wide Fisheye Camera
UIST '20· Full-Body Interaction & Embodied Input +2
Based on Jaccard similarity of research subtopics & professions (≥60%)