AI/ ai · benchmarks · vision-language-models · simulation

Benchmark Shows AI Agents Struggle to Audit 3D Worlds

A new benchmark tests whether vision-language models can both explore virtual environments and spot anomalies like floating objects, and most fail badly.

A new benchmark shows the AI agents meant to patrol virtual worlds still can't tell a floating couch from a normal one.

Researchers built WorldAuditBench, a test suite of 213 anomaly tasks spread across 13 environments made with Unreal Engine 5 and Three.js. The anomalies fall into five families: things like floating objects, walls you can walk through, and objects that don't match their surroundings. The team ran five frontier multimodal models through two setups: one where a vision-language-action model explores first and a vision-language model judges afterward, and one where a single VLM agent handles both exploring and spotting glitches at once. Under a fixed exploration budget, the models' success rates ranged from just 6.6% to 42.3%. Human testers hit 83.4%.

That gap matters because interactive 3D worlds are becoming a standard tool for training and evaluating embodied AI, from robotics simulators to game-based agents. If a model can't reliably tell a broken environment from a working one, it can't be trusted to flag errors before they corrupt downstream training data or mislead whatever system learns inside that world. The benchmark also isolates a specific weakness: these models struggle to combine looking, meaning visual reasoning, with doing, meaning navigating to verify a hunch, which is exactly the coupling embodied AI is supposed to be good at.

Call it a reality check for the idea that AI already understands 3D space. The best model here still missed more than half the planted glitches, and most missed the vast majority. Simulated worlds are only as useful as the ability to detect when they're broken, and right now, that job still belongs to humans.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →