[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-compression-trick-lets-coding-agents-share-gpus-better":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8227,"new-compression-trick-lets-coding-agents-share-gpus-better","New Compression Trick Lets Coding Agents Share GPUs Better","A new compression method lets AI coding agents serve 1.9 times more concurrent requests per GPU, trading a modest accuracy dip for far less memory use.","A new training trick lets AI coding agents forget most of what they've seen and still fix almost as many bugs.\n\nThe method, called LOHA (Latent Observations, Hard Actions), compresses older tool outputs, like error logs, file diffs, and test results, into compact soft tokens while keeping the agent's own reasoning and the last few observations as plain text. A companion training method, Anchored Context Distillation, teaches the agent to read that compressed history without drifting from how the original, uncompressed model behaves. On SWE-bench Verified, a benchmark built from real GitHub issues, this cut context size per call by 43% for a Qwen3-4B agent and 57% for a fine-tuned SWE-Master-4B-RL agent. Resolve rates dipped only slightly, from 14.5% to 12.1% for Qwen3 and from 27.5% to 21.8% for SWE-Master.\n\nThe advantage grows when memory is tight. Under a 32K-token limit, the compressed Qwen3 agent resolved 21.1% of a test subset versus 11.1% for the same agent running on full text, and in concurrent single-GPU serving it handled 1.9 times the request volume of its full-text counterpart. That is a direct answer to the real bottleneck in scaling coding agents: not raw intelligence, but how many can fit on the same hardware at once.\n\nIt is worth noting every compressed configuration still resolved fewer issues than its uncompressed twin, and the results so far cover only two 4B-parameter models on one benchmark, a preprint, not a shipped product.","[\"ai\",\"dev-tools\",\"llm-agents\",\"research\"]","2026-09-28T04:00:00.000Z","2026-09-28T17:48:56.348Z","2026-09-28T17:49:02.974Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix the dek: source reports 1.9x GPU throughput in concurrent serving, not that individual agents 'run nearly twice as fast' — restate as processing capacity, not speed, to avoid overstating the result.","resolved","ai",[30,32,33,34],"dev-tools","llm-agents","research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.31430",0,{"sections":41},[42,45,49,54,59,64,68,73,78,82,87,92,96,101],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",4844,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",762,{"name":50,"slug":51,"count":52,"latest_published_at":53},"Policy","policy",399,"2026-09-27T18:39:02.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",265,"2026-09-28T14:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",189,"2026-09-28T10:52:40.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":18},"Science","science",151,{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",135,"2026-09-26T14:30:00.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":79,"slug":32,"count":80,"latest_published_at":81},"Dev Tools",84,"2026-09-26T04:20:58.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",76,"2026-09-25T18:33:59.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":93,"slug":94,"count":90,"latest_published_at":95},"General","general","2026-09-26T17:02:42.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"Reviews","reviews",30,"2026-09-24T20:07:31.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]