AI/ autonomous driving · ai safety · interpretability · self-driving cars

Self-Driving AI Failure Monitors Don't Need Internal Access

A new audit of autonomous-driving failure monitors finds that peeking at a model's internal representations barely outperforms just watching its outputs.

Turns out you don't need to read a self-driving car's mind to guess when it's about to mess up.

Researchers audited a popular runtime-monitoring approach for autonomous vehicles: reading a model's internal representations to flag failures before they happen. They tested it on two systems, LaneSegNet, which builds vectorized road maps online, and VAD, an end-to-end driving planner. A latent-space probe trained on each model's internals could flag frame-level errors reasonably well, reaching an AUROC of 0.780 for spotting high mapping error in LaneSegNet and 0.868 for spotting bad trajectory predictions in VAD. But a simpler baseline that only looked at each system's visible outputs, the map predictions themselves for LaneSegNet, or the ego state, driving command, and planned trajectory for VAD, matched or beat the internal-access version, hitting 0.825 and 0.924 respectively. Stacking the latent features on top of those output-only baselines produced no statistically resolved improvement.

That result challenges a common assumption in AI safety circles: that peering inside a model reveals failure signals its outputs alone don't. If the outputs a system already produces carry nearly all the useful signal, then some of the complexity going into interpretability-based safety monitors may be buying little beyond extra compute and engineering overhead. For an industry racing to certify autonomous-driving stacks as safe, that is a cheap sanity check worth running before building elaborate internal-state monitors.

The researchers aren't writing off latent-space monitoring entirely. They are releasing an evaluation protocol and their failure labels so the next paper claiming internal access helps has to actually prove it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →