[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-a-math-framework-for-telling-real-ai-gains-from-lucky-guesses":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},10834,"a-math-framework-for-telling-real-ai-gains-from-lucky-guesses","A Math Framework for Telling Real AI Gains From Lucky Guesses","A new theoretical framework shows letting an AI vote on its own answers is provably sound, but letting it cherry-pick its best guess is not.","Agentic AI systems that seem to get smarter don't always get more correct - and a new paper tries to prove exactly when they do.\n\nResearchers propose a formal framework that treats an AI agent's verification process as a bounded computational \"stage\" with a strict terminal checker, then use complexity theory to separate different ways a system's score can rise: searching longer, getting more outside support, or changing how it generates and checks its own answers. They show that majority-vote amplification - running a system many times and taking the consensus answer - provably preserves correctness. But letting a system simply pick its single best-looking answer among many random attempts does not: it can let wrong outputs slip through disguised as verified. They also look at recursive self-improvement, where a system edits its own code, and prove that under a fixed, already-sound checking process, self-modification keeps the system within the same verification guarantees it started with rather than unlocking new provable capability.\n\nThat distinction cuts against a familiar AI marketing move: pointing to a jump in benchmark scores as proof of a genuinely smarter model, when the jump may just reflect more attempts and luckier picks. The paper's test family built around quota-enforced search makes this concrete, showing search success rates can climb by huge ratios with zero actual gain in which problems get correctly solved.\n\nFor anyone tracking the self-improving-agent hype cycle, the useful question isn't how much the number moved. It's what, exactly, was verified.","[\"agentic-ai\",\"ai-verification\",\"self-improvement\",\"complexity-theory\"]","2026-10-09T04:00:00.000Z","2026-10-09T17:29:08.748Z","2026-10-09T17:29:12.530Z","published",null,[],"ai",[26,27,28,29],"agentic-ai","ai-verification","self-improvement","complexity-theory",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.10611",0,{"sections":36},[37,40,44,49,54,59,63,68,73,78,83,88,93,98],{"name":38,"slug":24,"count":39,"latest_published_at":18},"AI",6606,{"name":41,"slug":42,"count":43,"latest_published_at":18},"Security","security",926,{"name":45,"slug":46,"count":47,"latest_published_at":48},"Policy","policy",486,"2026-10-08T22:40:11.000Z",{"name":50,"slug":51,"count":52,"latest_published_at":53},"Deals","deals",474,"2026-10-08T22:00:00.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Hardware","hardware",229,"2026-10-08T20:47:10.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":18},"Science","science",192,{"name":64,"slug":65,"count":66,"latest_published_at":67},"Consumer Tech","consumer-tech",181,"2026-10-08T23:26:35.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Startups","startups",117,"2026-10-08T16:45:00.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",114,"2026-10-08T17:57:01.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Dev Tools","dev-tools",105,"2026-10-07T16:59:11.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"General","general",66,"2026-10-09T04:46:11.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Gaming","gaming",58,"2026-10-08T20:08:45.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"Reviews","reviews",34,"2026-10-08T14:00:22.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"How-To","how-to",8,"2026-10-05T09:00:00.000Z"]