A new benchmark says today's self-driving planners still misjudge the rare, dangerous moments that matter most.
According to an arXiv paper titled ExceptionDrive (arXiv:2609.37871v1, posted September 30, 2026), researchers built a benchmark that inserts hazards into real nuScenes driving scenes, using vision-language models to screen and edit footage while preserving the original context. The benchmark spans 21 tasks across six safety categories, each defining a hazard zone, a required clearance margin, and an acceptable response. Because altering a scene invalidates the original recorded path, the authors designed a reference-free scoring system built around four metrics: Unsafe Rate, Hazard Clearance Compliance, Hazard Proximity Response, and Counterfactual Trajectory Shift, which together track how much a planner's path intrudes on or respects a danger zone. Testing seven planners, the paper's authors report that most of them frequently drove into hazard regions or left insufficient clearance.
This is the paper's own finding, not an independent audit, but it targets a real gap: benchmarks built from routine footage say little about how a planner behaves in the rare moment that actually counts. The authors also describe a Reminder Agent that identifies hazard type and suggests a high-level strategy to a separate vision-language decision model, without steering the vehicle itself, and report that in zero-shot tests this improved strategy accuracy and reduced under-warning.
Corner cases are the part of self-driving demos that companies tend to skip; a benchmark built specifically to surface them is a more honest test than another lap around a sunny closed course.