A small add-on module is making AI video-prediction systems more honest about the motion they actually track.
The system, called LPA-CWM, tackles a known weak spot in counterfactual world models (CWM), which infer object motion by comparing a video predictor's normal forecast against one where researchers nudge the scene. Older versions averaged every candidate prediction equally, with no sense of which ones made physical sense. The new Learned Physical Adjudicator is a 3.0-million-parameter module, trained on synthetic MOVi-F video data, that instead learns to weigh candidates based on physical plausibility, then runs a second targeted check to fill in gaps. Tested on DAVIS, Kinetics, and RoboTAP, it posted relative gains of 18.1% to 60.0% on the researchers' own completeness metric and also improved accuracy on an existing benchmark, TAP-Vid First.
The more interesting contribution might be the yardstick itself. The team built a new evaluation protocol, Completeness-aware Motion Correspondence, specifically because standard benchmarks can show low localization error while quietly missing entire chunks of an object's trajectory. That's a real problem for any system feeding motion predictions into robot planning or physical simulation, where a plausible-looking but incomplete trajectory can fail silently instead of obviously.
This is a research paper, not a product, and it's validated on curated benchmarks like DAVIS and RoboTAP rather than messy real-world footage. Still, it's a useful reminder that in motion AI, the metric you grade on shapes the blind spots you don't notice.