A new self-driving AI system just beat the human baseline on a major planning benchmark - not by dreaming up more routes, but by getting smarter about which route to trust.
The framework, called iDriveVLA, tackles what researchers describe as a generation-evaluation gap in multi-modal driving planners: models already produce a strong candidate trajectory most of the time, but often fail to pick it out of the pile. iDriveVLA adds a Safety-aware Scorer to judge trajectory quality and risk, plus a vision-language-model-guided Modulator that adjusts evaluation criteria based on the scene. Training happens in three stages - candidate imitation, candidate space refinement, and ranking alignment - meant to keep the evaluator in sync with an oracle standard. On the NAVSIM v1 leaderboard, it scored 94.95 PDMS, edging past the human-expert reference.
This matters because it reframes where progress in autonomous driving actually comes from. Instead of the industry's usual push toward generating more diverse trajectory options, this work suggests the real bottleneck is judgment - picking the right one under ambiguity. If that holds up outside this benchmark, it could redirect research effort toward evaluation architecture rather than ever-larger candidate generators.
Worth remembering: a leaderboard score is not a road test. Beating a "human-expert reference" on NAVSIM means winning a simulated scoring metric, not surviving a merge onto a real highway.