[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-oscar-framework-makes-small-open-weight-llms-nail-optimization-code":10,"sections":34},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":29,"feedback":33,"feedback_at":22,"cost_usd":33,"total_tokens":33},9490,"oscar-framework-makes-small-open-weight-llms-nail-optimization-code","OSCAR Framework Makes Small Open-Weight LLMs Nail Optimization Code","A new research framework routes optimization-coding tasks across cheap local models and still beats pricier AI coding tools on accuracy and cost.","A new academic framework lets two small, locally-run AI models solve business optimization problems about as well as pricier commercial ones, for a fraction of the cost.\n\nThe system, called OSCAR, tackles a specific failure mode: large language models can translate a plain-English business problem into optimization code, but that code sometimes misrepresents the actual constraints. A solver then confidently returns the optimal answer to the wrong problem. OSCAR's answer is to pair a Coder with a Reviewer and an offline Simulator, the last of which is checked against labeled examples of feasible and infeasible decisions so it can catch misreadings and keep searching for a better plan even after a feasible one turns up. The framework also decides, attempt by attempt, which LLM to call next and when to stop, escalating from cheap models to pricier ones only as needed. Across five benchmark problems, OSCAR hit 95 to 100 percent accuracy using two small open-weight models that each run on a single GPU - models whose single-attempt accuracy on their own averaged just 29 and 48 percent. Over five runs per problem, Codex and Claude Code cost 3.1 and 5.8 times more in tokens for the same work.\n\nThe real story here is the economics, not the accuracy number. Instead of chasing the biggest model for every task, OSCAR treats model choice as a cost-allocation problem and lets cheap, local models do most of the work while a verifier, not a bigger model, catches their mistakes. That's a different bet than most agent frameworks make, and it matters for any company running the same optimization task repeatedly, where confidentiality or GPU budgets rule out sending everything to a frontier model's API.\n\nThe catch: OSCAR only works because firms are expected to maintain labeled examples of good and bad decisions to keep the models honest. That's unglamorous homework most companies haven't done, and five benchmark problems are a long way from proving this scales to messier, real-world operating rules.","[\"ai\",\"optimization-modeling\",\"llm-routing\",\"open-weight-models\"]","2026-10-02T04:00:00.000Z","2026-10-02T21:42:21.211Z","2026-10-02T21:42:27.512Z","published",null,[],"ai",[24,26,27,28],"optimization-modeling","llm-routing","open-weight-models",[30],{"name":31,"url":32},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.00912",0,{"sections":35},[36,39,43,47,52,57,61,66,71,76,81,86,91,96],{"name":37,"slug":24,"count":38,"latest_published_at":18},"AI",5858,{"name":40,"slug":41,"count":42,"latest_published_at":18},"Security","security",833,{"name":44,"slug":45,"count":46,"latest_published_at":18},"Policy","policy",438,{"name":48,"slug":49,"count":50,"latest_published_at":51},"Deals","deals",317,"2026-10-01T22:00:00.000Z",{"name":53,"slug":54,"count":55,"latest_published_at":56},"Hardware","hardware",198,"2026-10-01T17:38:48.000Z",{"name":58,"slug":59,"count":60,"latest_published_at":18},"Science","science",171,{"name":62,"slug":63,"count":64,"latest_published_at":65},"Consumer Tech","consumer-tech",155,"2026-10-01T19:54:10.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Dev Tools","dev-tools",96,"2026-10-01T16:57:03.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Software","software",93,"2026-09-30T21:41:11.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Startups","startups",90,"2026-10-01T21:55:22.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]