[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-give-ai-coding-agents-runnable-checks-not-just-rules":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},9456,"researchers-give-ai-coding-agents-runnable-checks-not-just-rules","Researchers Give AI Coding Agents Runnable Checks, Not Just Rules","A new benchmark gave coding agents callable checks instead of written rules, and the advantage shrank or vanished once sample sizes grew.","A new benchmark asks a simple question: do AI coding agents patch broken scientific code better when handed a runnable check instead of a paragraph of rules.\n\nThe project, called Rules to Tools, pairs matched groups of agents repairing code from the SciCode benchmark, which packages equations, boundary conditions, and output requirements from real scientific software. Each group starts from identical code, the same model, and the same budget. One group gets the requirements as written text; the other gets a callable tool that actually runs the check. Across the two task-ID cohorts tested, the tool group completed repairs on 29 of 30 problems versus 26 of 30 for the text group, though the per-task picture was messier: three task IDs favored tools, one favored text, and eleven ended in a tie.\n\nZoom into the eight-task subset where the researchers ran their one statistical test, and the gap narrows to 15 of 16 versus 13 of 16, with a bootstrapped 95% confidence interval of -12.5 to 43.75 percentage points for that subset, a range wide enough to include \"no difference at all.\" On a larger shared-definition set of SciCode tasks, both groups tied at 13 of 24.\n\nThe more telling split shows up when the starting code doesn't match what a model likely memorized from training: on five such tasks, the tool group hit 7 of 10 versus 3 of 10 for text, suggesting runnable checks matter most when agents can't just pattern-match an answer. In a separate PDE comparison, checks also cut reported model output by 31.2% while matching text's accuracy, a real cost saving even where the correctness gap disappears.\n\nTranslate the small samples and overlapping intervals honestly: this reads as a plausible argument for executable specs over prose ones, not proof that handing an agent a tool beats handing it instructions.","[\"ai agents\",\"benchmarks\",\"scientific computing\",\"llm research\"]","2026-10-02T04:00:00.000Z","2026-10-02T19:56:29.224Z","2026-10-02T19:56:35.625Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The -12.5 to 43.75 percentage-point confidence interval is sourced to the eight-task-ID subset (13\u002F16 text vs 15\u002F16 tools), not to the full 29\u002F30-vs-26\u002F30 headline figure as the draft implies — attribute it to the correct subset or drop the 'headline result' framing.","resolved","ai",[32,33,34,35],"ai agents","benchmarks","scientific computing","llm research",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.00313",0,{"sections":42},[43,46,50,55,60,65,70,75,80,85,90,95,100,105],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",5764,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",831,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",437,"2026-10-01T18:10:00.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",317,"2026-10-01T22:00:00.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",198,"2026-10-01T17:38:48.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Science","science",168,"2026-10-01T18:35:55.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Consumer Tech","consumer-tech",155,"2026-10-01T19:54:10.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Dev Tools","dev-tools",96,"2026-10-01T16:57:03.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Software","software",93,"2026-09-30T21:41:11.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Startups","startups",90,"2026-10-01T21:55:22.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]