[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-agent-harness-beats-pricier-model-for-15-not-574":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},5161,"ai-agent-harness-beats-pricier-model-for-15-not-574","AI Agent Harness Beats Pricier Model For $15, Not $574","A new open-source runtime called StateM pushed a $15 AI agent run past the score of a $574 run using a pricier frontier model on a coding benchmark.","A new agent runtime called StateM gets language models to solve more of a demanding coding benchmark by changing how they manage state, not by changing the model.\n\nStateM organizes long agent runs around durable states, phase-local context, checked transitions, and versioned procedures instead of leaving each step to the model's own memory. On Terminal-Bench 2.1, it pushes GPT-5.5 xhigh's score to 92.1%, ahead of the 91.9% posted by the pricier GPT-5.6 Sol Ultra reference run, and the same runbook transferred to GPT-5.6 without modification. With GPT-5.6 Sol xhigh, StateM reaches 95.3% raw accuracy across 445 trials and clears all 89 tasks at least once. A similar frozen setup lifts GPT-5.6 Luna from 76.7% to 85.4%, and under $38 of tuning raises DeepSeek-V4 Flash from 82.7% to 88.1%.\n\nThe number that matters more than the accuracy bump is cost. StateM's winning run cost about $15 in API usage, versus $574.68 for the GPT-5.6 Sol Ultra reference it beat. That gap suggests the scaffolding around a model can matter as much as which model you buy, especially for teams priced out of frontier-tier agent runs.\n\nOn a benchmark built to separate models that can actually plan long tasks from ones that only look like they can, a cheaper model with better bookkeeping just outscored the expensive one.","[\"ai\",\"ai-agents\",\"benchmarks\",\"open-source\"]","2026-08-18T04:00:00.000Z","2026-08-18T07:41:21.489Z","2026-08-18T07:41:33.323Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The headline claims StateM only 'nearly matches' pricier models, but the body and closing line state it actually outscores them (92.1% beating GPT-5.6 Sol Ultra's 91.9%, and the $15 run outscoring the $574.68 reference run) — rewrite the headline\u002Fdek to reflect that the cheap harness beat the pricier baseline, not merely approached it.","resolved","ai",[30,32,33,34],"ai-agents","benchmarks","open-source",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.15089",0,{"sections":41},[42,46,50,55,60,65,70,75,80,84,89,94,99,104],{"name":43,"slug":30,"count":44,"latest_published_at":45},"AI",3293,"2026-08-20T04:00:00.000Z",{"name":47,"slug":48,"count":49,"latest_published_at":45},"Security","security",435,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":18},"Dev Tools","dev-tools",69,{"name":85,"slug":86,"count":87,"latest_published_at":88},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]