Concept bottleneck models (CBMs) force a neural network to predict human-readable concepts - like "has wings" or "is furry" - before making a final classification. The idea is that this extra step makes a model both interpretable and, researchers hoped, more robust to attacks. A new arXiv paper tests that assumption directly.
The authors built a generator-based evaluation framework to compare standard classifiers against CBMs under two kinds of perturbations: continuous geometric shifts in the model's latent space, and discrete semantic edits to the concepts themselves. They measured robustness both empirically, through prediction and concept-level sensitivity metrics, and certifiably, using randomized smoothing. The result reconciles years of contradictory findings in prior CBM research: interpretability doesn't uniformly make models more robust. It just moves the sensitivity somewhere else, with the effect depending heavily on the type of perturbation, how similar the target classes are, and how large the concept vocabulary is.
That distinction matters for anyone deploying CBMs in supposedly safety-critical or compliance-driven contexts, like medical imaging or hiring tools, where "interpretable" is sometimes treated as a proxy for "trustworthy." This paper says that's a category error: robustness and interpretability are separate engineering problems that need separate testing, not a package deal.
It's a useful corrective to the explainable-AI hype cycle, which has long assumed that opening the black box automatically makes it safer. Turns out you can see inside the box and still get burned - just by a different flame.