DancingBox: A Lightweight MoCap System for Character Animation from Physical Proxies
Honorable MentionAuthors
Paper Title
DancingBox: A Lightweight MoCap System for Character Animation from Physical Proxies
Publication Info
- Topic area: Lightweight motion capture system for character animation using everyday objects.
- Keywords: Motion capture, character animation, generative models, bounding boxes, physical proxies, puppetry, computer vision, motion diffusion, user study, tangible interfaces.
Background and Problem
- Problem / challenge: Traditional 3D character animation requires either expert use of complex software or expensive motion capture systems. Existing puppetry-based systems often rely on custom hardware or are limited to specific motion types.
- Significance: Making motion capture accessible to novices can democratize animation creation, enabling creative expression without specialized skills or equipment.
- Motivation and related work: Prior work has explored physical proxies and generative models for animation but is constrained by hardware requirements, limited motion types, or lack of generalization. This paper addresses these gaps by enabling motion capture with everyday objects and a single webcam.
Solution
- Proposed approach: DancingBox, a vision-based system that captures approximate object motions and refines them into realistic character animations using bounding-box representations and generative motion models.
- Novelty:
- Introduces a lightweight motion capture system using a single webcam and everyday objects.
- Employs bounding boxes as an intermediate representation to bridge coarse proxy motions and realistic character animations.
- Leverages generative motion models conditioned on bounding boxes to produce realistic animations.
- Synthesizes training data by converting motion capture datasets into proxy animations for training.
- Procedure and key techniques:
- Capture video of user manipulating a physical proxy with a single webcam.
- Segment and track proxy parts using vision foundation models (π3, SAM2, CoTracker).
- Represent proxy motions with 3D bounding boxes.
- Use a box motion encoder and a motion diffusion model (MDM) to generate realistic character animations.
- Evaluate the system through replication and creative tasks in a user study.
Results
- Concrete findings:
- System runtime: ~2 minutes 40 seconds on an NVIDIA 4090 GPU.
- User study showed participants could replicate motions in ~10–12 minutes, compared to an estimated one full day using traditional tools.
- Generated motions were rated as realistic and aligned with user intent, with a 92% success rate in replication tasks.
- Advantage over baselines:
- Supports arbitrary objects as proxies with a single webcam.
- Produces high-quality motion compared to prior systems with limited motion types or specific hardware.
- Requires no markers, suits, or calibration.
- Experiments / evaluation:
- User study with 9 participants performing replication and creative tasks.
- Evaluation metrics included task completion time, motion realism, user intent alignment, and user satisfaction.
- Participants used diverse proxies, including plush toys, bananas, and articulated puppets.
- Limitations and future work:
- Sensitive to occlusion due to monocular input.
- Non-real-time processing.
- Limited to human-like motions due to reliance on human motion datasets.
- Future work includes multi-character interaction, support for non-human characters, and additional user-requested features like speed control and integration with game engines.
Summary
DancingBox is a lightweight, vision-based motion capture system that enables novices to create realistic character animations using everyday objects and a single webcam. By leveraging bounding-box representations and generative motion models, the system transforms coarse proxy motions into detailed animations. A user study demonstrated its usability, effectiveness, and creative potential, with participants successfully replicating and designing diverse motions. While the system is limited by occlusion sensitivity and non-real-time processing, it offers significant advantages in accessibility and flexibility compared to traditional tools. Future work aims to extend its capabilities to multi-character interactions, non-human motions, and user-requested features.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)