A new computer vision paper breaks 3D body tracking into two separate problems and solves each one differently.
Researchers describe FactorizedHMR, a two-stage system for human mesh recovery, the task of estimating a 3D body's shape and pose from video even when a camera only sees part of a person. The first stage uses a deterministic regression model to lock down the torso and the body's root position, the part cameras usually capture well; a second, probabilistic stage then fills in the more uncertain limbs, like arms and legs, especially when they're occluded. To train the system, the team built a synthetic data pipeline pairing images, camera angles, and motion data across many viewpoints. On standard benchmarks, FactorizedHMR matched strong existing methods overall, with its clearest edge in two specific cases: heavy occlusion, and world-space tracking, where small errors tend to drift over time.
Most body-tracking systems treat the whole skeleton as one estimation problem, even though some joints are far more predictable than others. FactorizedHMR's bet is that separating a confident anchor from an uncertain, probabilistic reconstruction models that uncertainty more honestly than averaging it away. That distinction matters for motion capture, sports analytics, and AR avatars, where a flailing limb estimate breaks the illusion faster than a slightly-off torso.
The gains are real but narrow: competitive, not dominant, and concentrated in occlusion-heavy and drift-prone scenarios rather than across the board. That's normal for a benchmark paper, but worth remembering before anyone calls this a new standard.