AI/ ai · benchmarking · computer-vision · physics-simulation

AI Agents Ace Static Scenes, Fumble Dynamic Physics Benchmark

A new benchmark finds leading AI models can turn static scenes into code but struggle to simulate fluid, fracture, and deformation physics over time.

A new benchmark suggests today's best AI models can fake a photo but can't fake physics.

Researchers released 4DCodeBench, a benchmark testing whether AI agents can watch a video of a dynamic scene and write executable code that reproduces it, essentially reverse-engineering the motion, not just the picture. The test set combines real-world videos with synthetic scenes built around specific physical phenomena, including objects deforming, fluids flowing, and materials fracturing. To pass, an agent has to infer the scene's underlying structure, then implement something like a physics simulation that plays it back correctly instead of rendering a static 3D model. The researchers benchmarked several frontier models and found a clear gap between describing a scene and actually simulating it.

Static 3D reconstruction, turning a single image into a 3D model, has become routine for recent AI systems. This benchmark argues the harder problem is temporal: predicting how materials behave over time, which matters for robotics and for generated video that needs to look physically plausible rather than merely pretty. The finding suggests current models are good at mimicking appearance but still bad at reasoning about cause and effect.

It is the same lesson video-generation models keep relearning: nailing how something looks in one frame says nothing about whether a system understands how it would actually fall apart.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →