AI/ ai · llm-evaluation · context-compression · ai-agents

How AI context compression benchmarks hid failed runs

A new AI evaluation case study shows a benchmark quietly dropped incomplete runs, hiding how outcomes would shift if every boundary case counted.

AI agent benchmarks that look like a clean split can be quietly incomplete.

A case study of ReVerPi, a Pi framework extension that swaps old tool outputs for compressed, addressable excerpts, ran an 86-run campaign generating 641 model requests to compare full-context and projected-context runs. Fifteen matched pairs completed, with each arm succeeding on 12 of 15 tasks - a tie on paper. But twelve additional paired runs were cut short because the runner canceled the second arm whenever the first failed to finish; restoring all 27 boundary runs shows the projected method could land anywhere from 9 tasks worse to 1 task better than the full-context method. Of the fifteen completed pairs, only eleven had both arms succeed, and that smaller, fully-matched group is what the study actually used to compare resource use: projection cut total logical tokens by 25%, but the median pair still used 29% more tokens, with total follow-up requests rising from 35 to 55.

This is less a verdict on context projection than a warning about how evaluation scaffolding can hide failure: when a runner drops a companion test the moment its partner fails, real failures disappear from the scorecard instead of counting against it. The risk is concrete, not hypothetical - one dropped run had the projected-context agent burn twelve requests digging through archived text for an answer its full-context counterpart gave in three.

Context compression is marketed across agent frameworks as a free efficiency win; this study is a reminder to check what counts as "success," and how many runs got quietly excluded before the scorecard was drawn up.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →