[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-how-ai-context-compression-benchmarks-hid-failed-runs":10,"sections":45},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":35,"tags":36,"sources":40,"feedback":44,"feedback_at":22,"cost_usd":44,"total_tokens":44},8226,"how-ai-context-compression-benchmarks-hid-failed-runs","How AI context compression benchmarks hid failed runs","A new AI evaluation case study shows a benchmark quietly dropped incomplete runs, hiding how outcomes would shift if every boundary case counted.","AI agent benchmarks that look like a clean split can be quietly incomplete.\n\nA case study of ReVerPi, a Pi framework extension that swaps old tool outputs for compressed, addressable excerpts, ran an 86-run campaign generating 641 model requests to compare full-context and projected-context runs. Fifteen matched pairs completed, with each arm succeeding on 12 of 15 tasks - a tie on paper. But twelve additional paired runs were cut short because the runner canceled the second arm whenever the first failed to finish; restoring all 27 boundary runs shows the projected method could land anywhere from 9 tasks worse to 1 task better than the full-context method. Of the fifteen completed pairs, only eleven had both arms succeed, and that smaller, fully-matched group is what the study actually used to compare resource use: projection cut total logical tokens by 25%, but the median pair still used 29% more tokens, with total follow-up requests rising from 35 to 55.\n\nThis is less a verdict on context projection than a warning about how evaluation scaffolding can hide failure: when a runner drops a companion test the moment its partner fails, real failures disappear from the scorecard instead of counting against it. The risk is concrete, not hypothetical - one dropped run had the projected-context agent burn twelve requests digging through archived text for an answer its full-context counterpart gave in three.\n\nContext compression is marketed across agent frameworks as a free efficiency win; this study is a reminder to check what counts as \"success,\" and how many runs got quietly excluded before the scorecard was drawn up.","[\"ai\",\"llm-evaluation\",\"context-compression\",\"ai-agents\"]","2026-09-28T04:00:00.000Z","2026-09-28T17:45:23.234Z","2026-09-28T17:45:29.870Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Reconcile the run\u002Fpair counts against the source: the source says twelve boundary runs stopped (companions suppressed), but the article calls this '12 more pairs unresolved' and then cites 'restoring all 27 boundary runs' two sentences later — 12 pairs would mean 24 runs, not 27, so the boundary-run arithmetic needs to be fixed or clarified before publishing.","resolved",{"id":31,"reviewer":32,"round":33,"reason":34,"status":29},"publisher-r2","publisher",2,"The body is internally inconsistent about how many pairs completed — it first states 15 matched pairs finished with identical 12\u002F15 scores, then later refers to only 11 pairs having both modes run to completion for the token-usage comparison, without reconciling the discrepancy.","ai",[35,37,38,39],"llm-evaluation","context-compression","ai-agents",[41],{"name":42,"url":43},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.31381",0,{"sections":46},[47,50,54,59,64,69,73,78,83,88,93,98,102,107],{"name":48,"slug":35,"count":49,"latest_published_at":18},"AI",4844,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Security","security",762,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Policy","policy",399,"2026-09-27T18:39:02.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Deals","deals",265,"2026-09-28T14:00:00.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":68},"Hardware","hardware",189,"2026-09-28T10:52:40.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":18},"Science","science",151,{"name":74,"slug":75,"count":76,"latest_published_at":77},"Consumer Tech","consumer-tech",135,"2026-09-26T14:30:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Dev Tools","dev-tools",84,"2026-09-26T04:20:58.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Startups","startups",76,"2026-09-25T18:33:59.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":99,"slug":100,"count":96,"latest_published_at":101},"General","general","2026-09-26T17:02:42.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"Reviews","reviews",30,"2026-09-24T20:07:31.000Z",{"name":108,"slug":109,"count":110,"latest_published_at":111},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]