AI/ ai · interpretability · eu-ai-act · regulation

Study Finds AI Interpretability Evidence Flips 73% of the Time

Researchers found circuit-level AI interpretability evidence flips across 73% of defensible analysis choices, undermining EU AI Act filings.

The explanation a lab gives for how an AI model made a decision may depend more on which afternoon the analyst picked their settings than on the model itself.

Researchers ran a pre-registered test of circuit discovery, the leading method for explaining how neural networks reach specific outputs, on GPT-2 small performing a benchmark task called indirect object identification. They built a grid of seven analytic choices, each one already used in a published method, producing 15,840 possible specifications, of which 7,561 yielded an actual claim about which circuit did the work. The resulting explanation flipped across 73.2% of specification pairs, even though every individual choice was defensible on its own. The circuits behind those explanations barely overlapped, sharing a median of just 4% of components, and different runs' conclusions were statistically uncorrelated with each other.

This lands squarely on the EU AI Act, which requires providers of high-risk AI systems to file technical documentation explaining how their systems reach decisions, and circuit discovery is the leading candidate for producing that evidence. If two competent analysts using the same tool on the same system can produce contradictory official explanations, the paperwork becomes a coin flip dressed up as science. The researchers tried standardizing the single most influential setting, the evaluation metric, and the flip rate only dropped to 59.4%, still well above any plausible regulatory tolerance.

The study covers one small model and one narrow task, so its authors stop short of declaring the whole field broken. But the paper's real target, the tidy circuit diagrams regulators are expected to accept as evidence, look a lot less tidy under scrutiny.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →