AI/ coding-agents · formal-verification · ai-research · benchmarks

AI Coding Agents Get a Playbook for Verification Failures

A new verification toolkit nailed a perfect HumanEval score, but the real finding is that its audit step - not the proof - is what catches sabotaged code.

A new paper says the real bottleneck in trustworthy coding agents isn't writing proofs - it's catching the proofs that lie to you.

Researchers built a toolkit that pairs executable semantics in the K framework with procedures for writing specifications, fixing broken proofs, and auditing whether those proofs actually mean what they claim. Tested against HumanEval, a set of 164 Python programming tasks, the system hit 164 out of 164 successes, judged by AI audits that passed each program after at most two targeted repairs. To check whether that auditing step catches problems proofs alone would miss, the team built 12 pairs of near-identical clean and sabotaged code packages. Every package, clean or defective, passed its underlying K proof - but the audit step correctly flagged all 12 defective ones and cleared all 12 clean ones. A separate test, KleverBench, threw 31 programs with altered operator behavior at the system, and here results got messier: the reusable guidance did not clearly beat either strict acceptance rules or generic advice of equal length.

That gap matters more than the headline success rate. A perfect HumanEval score is a demo; a proof that silently passes a sabotaged package is the actual failure mode that erodes trust in "verified" AI-written code. The audit step, not the proof, is doing the real work of catching what looks correct but isn't - and the mixed KleverBench numbers suggest nobody has yet figured out how to make that auditing skill transfer cheaply.

As a reality check, the team also ran the method on Optimism's blockchain code, confirming expected pause-and-revert behavior for six operations under declared gas and semantics assumptions - a useful test, but a narrow sliver of what a genuinely verified coding agent will eventually need to handle.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →