A new study argues that how a video AI model learns matters more than how big it is when it comes to understanding the physical world.
Researchers ran four frontier self-supervised video models of matched size - V-JEPA 2.1, V-JEPA 2, VideoPrism, and VideoMAEv2 - through five stress tests: corrupted footage, fine-grained motion detection, occluded objects, reversed frame order, and general feature quality. The two V-JEPA models, which predict in an abstract "latent" space instead of reconstructing raw pixels, outperformed the others on every axis. When the team connected a frozen encoder to an action-conditioned planner in a simulated robot-arm task, only the latent-prediction models completed it reliably, kept working on degraded video, and transferred to a manipulation task the planner had never trained on.
Most video models still get graded on clean, curated clips, which says little about how they will behave inside a robot or a car navigating messy, partly blocked real-world scenes. This result suggests training method, not just scale or data volume, may decide whether a model reasons about cause and effect or just mimics the look of video. That reframes what counts as progress in a field currently obsessed with bigger datasets and bigger models.
Still, this is four models on two benchmarks and one simulated task, so treat it as a promising lead rather than a settled verdict until someone runs the same test on a real robot arm.