Dynamic gestures scatter more across people in speed and amplitude
Aliases: gesture speed variation · amplitude variation · DTW alignment · production variability
What it is
Across people, and across repetitions by one person, speed and amplitude of dynamic gestures fan out far more than static poses. The same “wave right” may be a wrist flick, an elbow sweep, or a half-metre arm path; it may finish in 200 ms or drag to a second. A pose template is a set of joint angles; individual difference sits mainly in hand size and range. A trajectory template is a path in time; alignment dominates whether it matches. Templates recorded from a few young adults at medium speed, then applied to children, older adults, or hurried operators, fail in a way that looks like “they cannot do it.” The distribution was never in the vocabulary or the recognizer.
Why it happens
Trajectory recognition aligns a test sequence to a template. Dynamic time warping can absorb some speed difference. Amplitude difference is a spatial scale: small moves drown in noise; large moves may leave the interaction volume. People also change amplitude with social setting and fatigue; it is not a stable personal constant. The mix of preparation and return varies too. A segmenter that assumes “the stroke is 60%” will cut a hasty stroke short. Pose recognition barely watches these time ratios. So the generalization bottleneck for a dynamic vocabulary is often not shape, but whose speed and whose amplitude were frozen as “a standard wave.” When sensor frame rate is low, fast productions are undersampled, and individual difference is amplified by the device.
Studying it
For each dynamic item, cover different ages or mobility, deliberately fast and slow, deliberately small and large, and before versus after fatigue. Report not only mean recognition rate but rate binned by speed and by amplitude, and which group the templates came from. Use leave-one-person-out so templates and tests are not the same hands. Run static poses on the same people as a control; pose accuracy usually sits flatter across individuals, while dynamic accuracy drops more steeply with speed and amplitude. Logging how much amplitude comes from wrist, elbow, and shoulder explains confusions that share a label but not a joint strategy.
Where it stops holding
A vocabulary of large, slow, ceremonial moves can crush individual difference by protocol, at the cost of slow learning and social visibility. Players in a game will exaggerate on purpose, so laboratory “natural” amplitude is never seen. Headset near-field tracking physically clamps amplitude; speed differences remain. People copying a teaching video converge on that video's tempo; without the demo, scatter returns. Static poses also scatter when hand size differs hugely (a child's hand versus an adult template), but that is scale, not a time path, and the correction is different.
Applying it
- Do not collect templates only from colleagues of one age at one speed. Deliberately include fast, slow, small, and large, and report recognition by bin.
- Let the recognizer normalize amplitude and time-warp speed, with caps: too small is noise, too large is out of bounds—not an infinite stretch onto the template.
- Teaching should state an acceptable speed and amplitude range (“a forearm sweep is enough; do not throw the whole arm”) and let people succeed at their own tempo in practice, not only on the video's beats.
Related
- Same group: C4.04.1 Static poses are shape; dynamic gestures are trajectory · C4.04.2 Static poses require holding and impose continuous muscle load
- Adjacent: C4.05 Three dimensions of gesture evaluation · C4.14 Gesture vocabulary size limits
- Search:
kinematic variability·gesture speed·amplitude normalization