AI/ ai · explainability · research · benchmarks

New Benchmark Tests Whether AI Explanations Survive Model Updates

A new study finds counterfactual explanations for AI decisions often fail to survive the very model changes they claim to withstand.

A new benchmark shows that "robust" AI explanations are often only robust to the one kind of model change their inventors happened to test.

Researchers built a unified evaluation protocol that tests six robust counterfactual explanation methods and two standard baselines against eight distinct types of model change, from small parameter tweaks to retraining on new data to full architecture swaps. They held the underlying data and generated explanations fixed across four tabular datasets, then measured how many predictions actually flipped under each change, along with coverage, validity, and proximity. Bounded parameter perturbations altered about 0.95% of test predictions on average, compared with 4.9% for bootstrap retraining, a roughly five-fold difference in how disruptive a "small" model update can be. Guarantees built for one type of change did not reliably transfer to others.

That's a problem because prior robustness claims were never really comparable, since each method had been graded against whichever single type of change its own authors chose to test. This benchmark puts them all through the same eight-way gauntlet instead of letting each one grade its own homework, and finds that relative performance and failure modes shift depending on which change you're testing against. One method, RobX, held up most consistently across the board, though the researchers note that its added stability can require larger interventions.

It's a reminder that "robust" in an explainability paper is a claim that needs a footnote specifying robust to what, not a blanket guarantee.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →