AI/ ai · benchmarks · embodied-ai · robotics

Benchmark Exposes How AI Agents Fake Physical Planning

A new causal-reasoning benchmark shows top AI models, including GPT-6-astra, mostly predict plausible words rather than track real cause and effect.

A new benchmark suggests today's embodied AI agents are better at sounding right than being right about physics.

Researchers built Causal-Plan-Bench, a diagnostic benchmark spanning four causal dimensions and checked through multiple rounds of verification, meant to separate models that track real cause-and-effect from ones that just pattern-match language. They also built Causal-Plan-1M, a million-scale dataset of causal reasoning traces pulled from first-person video through a four-stage annotation pipeline. Tested on the new benchmark, even GPT-6-astra, currently one of the stronger models available, scored just 43.04. Using their dataset to retrain the open Qwen3-VL-8B model into what they call Causal Planner pushed its score from 33.23 to 45.28, a 36.3 percent relative gain, and the improvement carried over to three other external benchmarks without extra tuning for them.

The result gives a concrete shape to a gap that's mostly stayed theoretical: a language model can describe a plan fluently while having no real grip on the physical consequences of each step. That distinction matters for anything meant to act in the real world, from warehouse robots to home assistants, where a plausible-sounding next move that ignores physics just means a dropped box or a jammed door. The researchers also report a causal-supervision scaling trend, meaning more of this training data keeps producing gains rather than plateauing right away.

Still, a purpose-built model trained on the researchers' own causal dataset topping out at 45 out of 100 says less about a breakthrough than about how far even specialized systems still have to go.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →