[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-benchmark-exposes-where-ai-code-reasoning-breaks-down":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},11029,"new-benchmark-exposes-where-ai-code-reasoning-breaks-down","New Benchmark Exposes Where AI Code Reasoning Breaks Down","A new benchmark finds that longer execution traces quietly flip hundreds of correct AI predictions to wrong ones, even with reasoning on.","A new benchmark shows AI models that ace short code-execution puzzles often stumble once the same program runs a little longer.\n\nResearchers built 400 test cases from 371 Python and C++ programs, extending the CRUXEval-style format that asks a model to predict a program's output just by reading the code and its input, no running it allowed. The new wrinkle: each program comes in a shorter-trace and longer-trace version, plus checkpoint tasks that ask the model to state the exact program state mid-loop and after the loop. Four model families were tested across seven settings, producing 11,151 gradable predictions out of 11,200 attempted. Turning on reasoning mode mattered a lot, lifting accuracy by 33.1 to 55.2 percentage points over non-reasoning settings on completed responses.\n\nEven the best setup shows the cracks: 93.0% accuracy on short-trace output prediction fell to 77.0% on the longer-trace version of the identical code, and dropped further to roughly 64% on the two checkpoint-state tasks. In 2,397 matched Python comparisons using the exact same source, swapping in the longer trace alone flipped 528 previously correct answers to wrong, with only 147 flipping back. The researchers note the swapped inputs change several things at once, so this doesn't prove models fail at tracking state specifically, just that something about longer traces trips them up.\n\nFinal-answer accuracy has long been the industry's go-to scoreboard for reasoning. This benchmark is a decent reminder that acing a short snippet and actually tracking what a program is doing are not the same skill.","[\"ai benchmarks\",\"code reasoning\",\"llm evaluation\",\"research\"]","2026-10-09T04:00:00.000Z","2026-10-10T02:50:20.067Z","2026-10-10T02:50:24.273Z","published",null,[],"ai",[26,27,28,29],"ai benchmarks","code reasoning","llm evaluation","research",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.11889",0,{"sections":36},[37,40,44,49,54,58,62,67,72,77,81,86,91,96],{"name":38,"slug":24,"count":39,"latest_published_at":18},"AI",6783,{"name":41,"slug":42,"count":43,"latest_published_at":18},"Security","security",932,{"name":45,"slug":46,"count":47,"latest_published_at":48},"Policy","policy",486,"2026-10-08T22:40:11.000Z",{"name":50,"slug":51,"count":52,"latest_published_at":53},"Deals","deals",474,"2026-10-08T22:00:00.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":18},"Hardware","hardware",232,{"name":59,"slug":60,"count":61,"latest_published_at":18},"Science","science",193,{"name":63,"slug":64,"count":65,"latest_published_at":66},"Consumer Tech","consumer-tech",181,"2026-10-08T23:26:35.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Startups","startups",117,"2026-10-08T16:45:00.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Software","software",114,"2026-10-08T17:57:01.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":18},"Dev Tools","dev-tools",106,{"name":82,"slug":83,"count":84,"latest_published_at":85},"General","general",66,"2026-10-09T04:46:11.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Gaming","gaming",58,"2026-10-08T20:08:45.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Reviews","reviews",34,"2026-10-08T14:00:22.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"How-To","how-to",8,"2026-10-05T09:00:00.000Z"]