[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-target-ais-silent-optimization-code-errors":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},5582,"researchers-target-ais-silent-optimization-code-errors","Researchers Target AI's Silent Optimization Code Errors","ReLoop checks AI-generated optimization code by testing its behavior, not just whether it runs, closing a gap that hit 90 percentage points on hard problems.","AI models can write optimization code that runs cleanly and still gets the problem wrong, and a new system called ReLoop is built to catch it.\n\nResearchers found that large language models translating natural-language requests into optimization code often produce programs that execute without errors and return solver-feasible answers that are nonetheless wrong, a feasibility-correctness gap that reached 90 percentage points on compositional problems. ReLoop attacks this with two mechanisms: structured generation, which splits code writing into four stages (understand, formalize, synthesize, verify) to stop mistakes at the source, and behavioral verification, which tests how a formulation responds to solver-based parameter perturbation, an external check that does not depend on the model reviewing its own work. Paired with a diagnostic recovery step, the combined system reports 100 percent executable code and consistent accuracy gains across three benchmarks. The team also released RetailOpt-190, 190 retail optimization scenarios designed to expose the multi-constraint interactions where LLMs most often fail.\n\nThe finding matters because optimization code is a plausible-looking output that is hard to audit. It compiles. It runs. A solver hands back a number. Nothing about that process tells you whether the model understood the problem it was asked to solve, which is exactly the failure mode ReLoop is designed to surface.\n\nOne number in the paper deserves a caveat: the largest single-benchmark gain, 8.5 points on RetailOpt-190, is credited to a model the paper calls Claude Opus 4.6, a designation that does not correspond to any Claude model Anthropic has publicly released, so that specific figure should be treated as unverified until the model is properly identified.","[\"ai\",\"llm\",\"dev-tools\",\"research\"]","2026-08-18T04:00:00.000Z","2026-08-19T02:20:05.739Z","2026-08-19T02:20:17.571Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Verify the model identifier 'Claude Opus 4.6' against Anthropic's actual published model lineup before publishing — it does not match any known real-world release and needs correction or sourcing before this can run.","resolved","ai",[30,32,33,34],"llm","dev-tools","research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2602.15983",0,{"sections":41},[42,46,50,55,60,65,70,75,80,83,88,93,98,103],{"name":43,"slug":30,"count":44,"latest_published_at":45},"AI",3293,"2026-08-20T04:00:00.000Z",{"name":47,"slug":48,"count":49,"latest_published_at":45},"Security","security",435,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":81,"slug":33,"count":82,"latest_published_at":18},"Dev Tools",69,{"name":84,"slug":85,"count":86,"latest_published_at":87},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]