[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-framework-audits-ai-written-lean-proofs-before-acceptance":10,"sections":49},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":39,"tags":40,"sources":44,"feedback":48,"feedback_at":22,"cost_usd":48,"total_tokens":48},9678,"framework-audits-ai-written-lean-proofs-before-acceptance","Framework Audits AI-Written Lean Proofs Before Acceptance","A new auditing layer catches AI coding agents that compile Lean proofs without actually proving the claimed statement, tested on 772 problems.","A new framework checks whether AI-written Lean proofs actually prove what they claim, not just that they compile.\n\nResearchers released FORALL-LEAN-AGENT, a review layer that sits between AI coding agents and the Lean proof assistant, checking that a compiled proof actually proves the intended statement under legitimate assumptions. It compares the proof's final statement to the target, audits which axioms it leaned on, and runs independent proof checking where supported, with every decision traceable to the specific proof artifact it evaluated. On a 100-task slice of VeriSoftBench, wrapping GPT-5.6 Sol in the framework pushed success from 93 to 100 while cutting average cost from $69 to $62 per task. Run against all 672 problems in PutnamBench, it accepted every one at an average of $4.72 each.\n\nA proof that compiles isn't the same as a proof that's correct and non-trivial - a permissive checker can let shortcuts or weaker axioms slide through undetected. By auditing axioms and verifying statements independently, this framework targets exactly the gap that lets flawed proofs pass as valid. That distinction matters more as companies lean on formal verification to vouch for safety-critical code and mathematical claims made by AI systems.\n\nStrong numbers on two curated benchmarks are still a long way from catching every way a probabilistic model can game a formal checker, and the real test will be how this holds up on messier problems than Putnam-style math.","[\"ai\",\"formal-verification\",\"lean-proofs\",\"benchmarks\"]","2026-10-02T04:00:00.000Z","2026-10-03T05:55:29.249Z","2026-10-03T05:55:33.229Z","published",null,[24,30,35],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Verify or drop the model name \"GPT-5.6 Sol\" — it does not match any known real model, and the article cannot cite an unverifiable model name as the basis for the headline's 93-to-100 figure.","resolved",{"id":31,"reviewer":32,"round":33,"reason":34,"status":29},"publisher-r2","publisher",2,"The body states the framework was tested on three benchmarks (VeriSoftBench, PutnamBench, and Lean Eval) but only reports results for the first two, leaving the Lean Eval outcome entirely missing.",{"id":36,"reviewer":26,"round":37,"reason":38,"status":29},"editor-r3",3,"The draft explicitly flags that Lean Eval success rate and cost are missing for both problems yet still runs with the claim that the framework was tested on three benchmarks — either report the actual Lean Eval results or cut that benchmark from the piece entirely before publishing.","ai",[39,41,42,43],"formal-verification","lean-proofs","benchmarks",[45],{"name":46,"url":47},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.00885",0,{"sections":50},[51,54,58,62,67,71,75,80,85,90,95,100,105,110],{"name":52,"slug":39,"count":53,"latest_published_at":18},"AI",5975,{"name":55,"slug":56,"count":57,"latest_published_at":18},"Security","security",842,{"name":59,"slug":60,"count":61,"latest_published_at":18},"Policy","policy",438,{"name":63,"slug":64,"count":65,"latest_published_at":66},"Deals","deals",317,"2026-10-01T22:00:00.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":18},"Hardware","hardware",199,{"name":72,"slug":73,"count":74,"latest_published_at":18},"Science","science",173,{"name":76,"slug":77,"count":78,"latest_published_at":79},"Consumer Tech","consumer-tech",155,"2026-10-01T19:54:10.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Dev Tools","dev-tools",96,"2026-10-01T16:57:03.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Software","software",93,"2026-09-30T21:41:11.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Startups","startups",90,"2026-10-01T21:55:22.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":111,"slug":112,"count":113,"latest_published_at":114},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]