Researchers have a new way to make self-driving cars plan safer routes without slowing down the decision loop that steers the car.
The approach, called EMPlan, splits trajectory planning into two steps. A lightweight module first proposes a handful of coarse candidate paths, called sparse anchors, then a second module refines those anchors into precise trajectories. Training happens in two stages: a standard pretraining phase, followed by reward-guided fine-tuning that uses rule-based reward signals and unpaired preference data to nudge the model toward safer choices. Because that fine-tuning happens during training rather than at inference time, it does not add extra computation when the car is actually driving. The method was tested on NAVSIM, a non-reactive benchmark used to compare planning systems.
This matters because the usual fixes for planning safety all have a catch. Copying human driving demonstrations inherits human blind spots. Scoring every candidate path against hand-written rules is accurate but computationally heavy. Preference-based training usually demands pairs of labeled examples, which are expensive to collect. EMPlan's unpaired approach loosens that last constraint while keeping the safety signal, which is the kind of unglamorous plumbing fix that tends to matter more than it sounds.
Still, this is a benchmark result on a non-reactive simulation, not a car on a real road. Plenty of planning methods look great against NAVSIM and never make it past a press release.