A new framework checks whether AI-written Lean proofs actually prove what they claim, not just that they compile.
Researchers released FORALL-LEAN-AGENT, a review layer that sits between AI coding agents and the Lean proof assistant, checking that a compiled proof actually proves the intended statement under legitimate assumptions. It compares the proof's final statement to the target, audits which axioms it leaned on, and runs independent proof checking where supported, with every decision traceable to the specific proof artifact it evaluated. On a 100-task slice of VeriSoftBench, wrapping GPT-5.6 Sol in the framework pushed success from 93 to 100 while cutting average cost from $69 to $62 per task. Run against all 672 problems in PutnamBench, it accepted every one at an average of $4.72 each.
A proof that compiles isn't the same as a proof that's correct and non-trivial - a permissive checker can let shortcuts or weaker axioms slide through undetected. By auditing axioms and verifying statements independently, this framework targets exactly the gap that lets flawed proofs pass as valid. That distinction matters more as companies lean on formal verification to vouch for safety-critical code and mathematical claims made by AI systems.
Strong numbers on two curated benchmarks are still a long way from catching every way a probabilistic model can game a formal checker, and the real test will be how this holds up on messier problems than Putnam-style math.