AI/ ai agents · benchmarks · spatial reasoning · minecraft

New Benchmark Finds Big Gaps in AI Spatial Navigation Skills

A new Minecraft-based benchmark shows leading AI models struggle to navigate recreated real-world landmarks, and open-weight models lag far behind the leaders.

AI agents still get lost inside recreations of Buckingham Palace and Midtown Manhattan.

Researchers built Mine Odyssey, a benchmark that rebuilds 30 real-world locations across 20 countries and regions inside Minecraft, covering 20 outdoor settings and 10 indoor ones, from rural Entrup to Santiago's Santa Lucia Hill. The 180 tasks each give an agent a plain-language instruction naming a sequence of waypoints, such as landmarks, buildings, or specific rooms, that testers manually checked for accessibility first. Completing a task means finding real entrances, opening doors, climbing stairs or ladders between levels, and noticing when a route has gone wrong and backtracking.

Eight models were put through the gauntlet, and the spread is the real headline here. GPT-6 Astra led with an 85.6% success rate, Claude Opus 5.5 followed at 73.9%, and DeepSeek-V4.1-Flash, the strongest open-weight model tested, finished at just 23.9%. That three-way gap matters for anyone shopping for a model to drive a delivery robot or a warehouse picker, since claims about spatial reasoning clearly mean very different things depending on which model is actually doing the walking.

Mine Odyssey tests bots in a block world, not a real parking lot, but a system that still can't reliably find the exit from Buckingham Palace is nowhere near ready to run your warehouse floor.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →