A new benchmark shows top AI reasoning models will confidently hand over a formula for a cause-and-effect question that has no valid answer.
Researchers built a tool called CERTID to test whether language models can tell when a causal effect actually can be calculated from the data they're given, and when it can't. The tool uses a proven algorithm to certify, ahead of time, which questions have no correct answer, then checks any formula a model produces against causal models where the real answer is already known. The team ran Gemini Flash, Gemini Pro, and GPT-5.5 through 1,200 of these certified test cases, covering causal graphs from 4 to 50 variables. The models answered correctly most of the time, but the rate at which they invented a false formula for an unanswerable question varied seventeen-fold between models, even on identical questions.
That gap matters because a wrong answer to an unanswerable question is nearly impossible to catch. If a model refuses to answer, that's an obvious miss. If it answers a question that provably has no solution, there's no dataset you can check it against to prove the formula wrong - it just looks like real math. The researchers also found models correctly judged whether a question was answerable 97 to 100 percent of the time on graphs built after the strongest model's training cutoff, meaning this isn't just the model echoing memorized problem sets.
Overall accuracy is the number most benchmarks report, and it's exactly the number this paper says can't be trusted on its own. A model can score well while still being happy to fabricate unfalsifiable statistics whenever the question is phrased like causal inference. For anyone using these models to reason about policy or experiment design, that's a more useful warning than another leaderboard score.