[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-study-finds-ai-coding-agents-fail-when-commands-get-reparsed":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},4952,"study-finds-ai-coding-agents-fail-when-commands-get-reparsed","Study Finds AI Coding Agents Fail When Commands Get Reparsed","A new benchmark shows that whether an AI agent's shell command succeeds says almost nothing about whether its output survives the pipe safely.","A new academic benchmark shows that grading an AI coding agent purely on whether its shell command worked can hide exactly how it's failing.\n\nResearchers built QuoteBench from 56 one-shot tasks drawn from 14 real incident families, testing agents that send Bash commands through code which serializes, wraps, and reparses the model's raw output before running it. They inserted one deliberately unescaped parser into that path and compared results with and without telling the model the boundary existed. Replaying the exact same generated reply through the flawed parser cut success rates by 55.4 to 73.2 percentage points across eight test configurations. Disclosing the parsing boundary to the model recovered 30.4 to 60.7 points in six of those configurations, and did almost nothing in the other two.\n\nThe headline finding is that a single matched-execution score, the number most benchmarks report, can mask this entirely. One configuration studied, which the paper labels GPT-5.6-sol - a designation specific to this study, not a publicly documented release we could independently verify - showed a near-flat overall gap of just -3.6 points. That number was quietly averaging out 64.3 points of damage from the broken parser against 60.7 points of recovery once the boundary was disclosed. Two identical-looking scores, two completely different agents underneath.\n\nThe paper also found that swapping the deployment plumbing, not the model itself, can flip which agent wins a head-to-head comparison in at least one clear case among 26 pairs tested. That's a quiet indictment of leaderboards that report a single number without saying how commands actually got from model to shell.","[\"ai agents\",\"llm evaluation\",\"benchmarks\",\"ai safety\"]","2026-08-14T04:00:00.000Z","2026-08-14T20:30:23.912Z","2026-08-14T20:30:35.748Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The model name \"GPT-5.6-sol\" is not a recognizable or plausible existing model designation — confirm it's the paper's actual terminology and flag it as such (or verify against the real model name) before publishing, since citing an unfamiliar version identifier without verification is a known readiness failure.","resolved","ai",[32,33,34,35],"ai agents","llm evaluation","benchmarks","ai safety",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.13547",0,{"sections":42},[43,47,51,56,61,66,71,76,81,86,91,96,101,106],{"name":44,"slug":30,"count":45,"latest_published_at":46},"AI",3293,"2026-08-20T04:00:00.000Z",{"name":48,"slug":49,"count":50,"latest_published_at":46},"Security","security",435,{"name":52,"slug":53,"count":54,"latest_published_at":55},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Dev Tools","dev-tools",69,"2026-08-18T04:00:00.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":107,"slug":108,"count":109,"latest_published_at":110},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]