Four frontier AI models each got a turn driving a real Toyota Corolla around a cone course, and only one reached the finish line.
Researchers built DrivingBench, a new benchmark that hands vision-language models a camera feed from the car and lets them issue direct steering and speed commands through tool calls, no simulator involved. The four models tested were OpenAI's GPT-6 Astra and GPT-5.6 Sol, Anthropic's Claude Fable 5.1, and xAI's Grok 4.6, each running inside the agentic scaffold its vendor normally ships for coding work: the two OpenAI models through Codex, Claude Fable 5.1 through Claude Code, and Grok 4.6 through Cursor. Those harnesses were repurposed as a tool-calling loop, not writing software, just reading camera frames and outputting steering and velocity commands, and the car kept rolling while each model was still "thinking," so slow inference meant stale commands got overwritten mid-maneuver. Each model got up to three attempts in a single conversation, and only GPT-6 Astra finished the course, doing so on its second try; no other attempt covered even half the layout.
The real finding isn't the leaderboard, it's that inference latency became part of the test itself: a model that reasons carefully but slowly effectively steers worse than one that acts fast and sloppy, because the car doesn't wait for it. That's a meaningfully different skill than the benchmarks these models usually top, writing code, passing exams, holding a conversation, and it's a reminder that being fluent in a coding harness says nothing about handling a continuous, physical, time-pressured task. Two of the four models did get visibly better across their three attempts when allowed to keep context from the failed runs, suggesting part of the bottleneck is learning the course rather than raw capability.
Call it a parking test for AI: cheap to run, hard to fake, and a useful antidote to assuming chatbot benchmarks translate into anything resembling real-world competence.