Code-generating AI models are just as confident when their code fails as when it works.
Researchers tested four open-source code generation models against three execution-based benchmarks, checking whether a model's own confidence score lined up with whether its code actually ran correctly. Mostly, it did not. Incorrect programs often scored confidence levels indistinguishable from correct ones, both at the level of whole programs and individual tokens. Instruction-tuned models made the problem worse in one specific way: they got more certain about their answers without getting more accurate. Standard fixes, like filtering out low-confidence outputs before accepting them, did not reliably close the gap either.
That matters because developers increasingly treat a confident-sounding AI coding tool as a signal to skip double-checking its work. This research argues that confidence is not a trustworthy stand-in for correctness, which is a problem as more teams lean on AI-generated code with lighter review. The researchers did find a silver lining: a model's internal hidden representations seem to encode correctness-related information that never makes it into the confidence score it actually reports.
In other words, the model may "know" it is wrong somewhere inside. It just is not telling you.