AI/ ai · computer-vision · robotics · research

New AI System Estimates an Object's Hidden Center of Mass

A research team built a video AI that estimates where an opaque object's mass is hidden, though it tends to overshoot the real target point.

A new system can estimate where an opaque object's weight is hidden just by watching a few seconds of video.

Researchers built STATERA, which takes a frozen, pretrained video model called V-JEPA and adds a small trainable module to track motion over time. Rather than reading pixel level shape or color, it watches how an object tumbles and infers where its center of mass sits inside. The team built a new benchmark called HiddenMass to test this: 50,000 simulated trajectories from the MuJoCo physics engine, plus 63 real-world video clips with center-of-mass locations measured by hand. In simulation, the best version of STATERA cut the normalized error in locating an object's center of mass from 41.7 percent, set by a baseline model called DINOv2, down to 25.2 percent.

The more telling result came when the model moved from simulation to real footage with zero retraining. A version trained with motion-phase information was the only one that consistently pushed its guesses toward the true hidden offset, though that same sensitivity caused it to overshoot the target and rack up slightly more raw positional error than just guessing an object's visible geometric center. That same version also got far better at capturing the physics of the motion itself, not just the final position, improving a separate accuracy metric from 2.6 percent to 41.0 percent. A version trained without phase information avoided the overshoot problem, but only by defaulting to statistically safe centroid guesses instead of real commitments.

It is a small, specific result, but it says something larger: a video model trained only to recognize what things look like can be coaxed into inferring physics it was never directly taught, and a confidently wrong guess can still beat a cautious shrug.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →