A new benchmark shows today's flashiest video-generation AI models can't roll a fair die - literally.
Researchers built CaliBench to test whether video world models, the systems that generate short video clips predicting how a scene might unfold, actually reproduce the correct odds of random physical events. Instead of judging video realism the way older benchmarks do, CaliBench scores clips against known probability distributions: a Galton board's binomial spread, a binary Bernoulli fork, uniform dice, cards and lottery draws, and a skewed European roulette color wheel. The team ran nine such scenes through six leading image-to-video models - WAN-2.7, SeeDance-2.0, HappyHorse-1.0, Veo 3.1, Runway Gen-4.5 and Cosmos3-Super - generating 32 clips per scene per model. They then measured both how often a model produced a readable outcome at all, a metric they call scorability, and how far its outcomes drifted from the true odds, or calibration, checking for statistical significance with a chi-squared test.
This matters because video world models are increasingly pitched as stand-ins for physics simulators - training tools for robotics, autonomous driving and game AI that need to reason about uncertain outcomes, not just pretty pixels. A model that quietly funnels every dice roll toward the same number, as Veo 3.1 did in testing, isn't simulating physics. It's memorizing a plausible-looking guess.
It's a useful reminder that a video clip looking real and a video clip being statistically honest are two different engineering problems - and going by these results, none of the six labs tested have solved the second one.