A new model watches ordinary video and calculates not just how a person is moving, but how hard they're pushing against the world around them.
The system, called PACT, comes from a paper posted to arXiv. It builds on an existing human-pose reconstruction model by adding learnable contact-force tokens and a temporal transformer that ties visual features to world-space motion. Joint prediction heads then refine pose estimates while estimating contacts and forces together, with physics-based supervision keeping the motion and force outputs consistent. Because labeled force data is scarce, the researchers built their own annotation pipeline combining contact labeling with physics-based optimization, and introduced a new benchmark, ForceWall, using climbing videos matched against real force-sensor readings.
Most pose-estimation systems handle motion, contact, and force as separate stages, which lets errors in one stage bleed into the next. PACT's pitch is that learning all three jointly avoids that compounding error and generalizes better to interactions the model never saw in training. If it holds up, that matters for robotics, biomechanics, and sports analysis, where understanding physical interaction from plain video - not motion-capture rigs - would be a real practical win.
Worth flagging: this is a single preprint, not yet peer reviewed, and the paper doesn't publish numeric head-to-head comparisons against named rival systems - so the state-of-the-art claim rests on the authors' own tables for now.