[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-letting-models-write-code-stops-reasoning-from-collapsing":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},9810,"letting-models-write-code-stops-reasoning-from-collapsing","Letting Models Write Code Stops Reasoning From Collapsing","A new benchmark finds model reasoning accuracy collapses on deep problems unless they can write and execute code instead of reasoning in plain text.","Large language models' reasoning accuracy collapses on deep problems unless they are allowed to write and run code instead of just thinking in words.\n\nResearchers built a benchmark around Boolean circuits over GF(2), a mathematical structure that can represent any computable function. The design strips out two things that make typical reasoning benchmarks hard to trust: the chance a model memorized the answer, and the ambiguity in how a problem can be broken into steps. Models solved these circuits step by step while researchers controlled circuit depth directly. When reasoning purely in natural language, next-step accuracy fell apart as depth increased, a pattern that held for both small and frontier models.\n\nThat matters because out-of-distribution reasoning, the ability to recombine learned rules to solve problems a model has not seen before, is the capability underpinning claims about general intelligence in AI systems. The benchmark suggests that capability is shakier than headline scores imply. When models were instead allowed to synthesize and execute code to solve the same circuits, the depth-related collapse did not happen, in either the small or frontier models tested.\n\nThe fix for brittle reasoning here is not more training data. It is a calculator. That is a humbling finding for systems marketed as reasoning engines.","[\"ai research\",\"benchmarks\",\"tool use\",\"reasoning\"]","2026-10-02T04:00:00.000Z","2026-10-03T11:33:37.350Z","2026-10-03T11:33:42.299Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Remove or qualify the closing line's invented specific that models need a Python interpreter 'to get past four or five logical steps' — the source doesn't give a step count, so this reads as a fabricated statistic; end on a closing line supported by what the benchmark actually showed.","resolved","ai",[32,33,34,35],"ai research","benchmarks","tool use","reasoning",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2602.21061",0,{"sections":42},[43,46,50,54,59,63,67,72,77,82,87,92,97,102],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",6058,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",848,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Policy","policy",439,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",317,"2026-10-01T22:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":18},"Hardware","hardware",199,{"name":64,"slug":65,"count":66,"latest_published_at":18},"Science","science",176,{"name":68,"slug":69,"count":70,"latest_published_at":71},"Consumer Tech","consumer-tech",155,"2026-10-01T19:54:10.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Dev Tools","dev-tools",96,"2026-10-01T16:57:03.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":81},"Software","software",93,"2026-09-30T21:41:11.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",90,"2026-10-01T21:55:22.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]