A new framework aims to stop interpretable AI models from faking their own explanations.
Concept bottleneck models, or CBMs, are built to be legible: before producing a final answer, they first predict human readable 'concepts' (say, 'has wings' or 'is metal') and base the prediction on those. Researchers found that CBMs often ignore the stated logic between concepts, or quietly leak information through supposedly separate steps, undermining the transparency they are built for. The new framework, called CREAM, lets engineers directly encode known relationships, like mutual exclusivity, hierarchy, or correlation, between concepts, plus how sparse each concept's influence on the final task should be. It also allows an optional side channel to cover gaps when the concept set is incomplete, paired with a new metric to measure how much that side channel is doing the real work.
This matters because concept bottleneck models are pitched for high stakes fields like medical diagnosis and credit decisions, where a prediction needs a real, checkable reason, not just a plausible sounding one. Earlier CBMs offered interpretability in name only if their internal concept logic did not match reality; CREAM's experiments show the architecture can hit black box level accuracy even with missing concepts, which is the usual excuse for abandoning interpretable designs altogether.
Of course, an optional side channel is also exactly the kind of escape hatch that could quietly become a black box by another name, and this is still lab bound research without a production deployment to test that claim.