AI agent benchmarks that look like a clean split can be quietly incomplete.
A case study of ReVerPi, a Pi framework extension that swaps old tool outputs for compressed, addressable excerpts, ran an 86-run campaign generating 641 model requests to compare full-context and projected-context runs. Fifteen matched pairs completed, with each arm succeeding on 12 of 15 tasks - a tie on paper. But twelve additional paired runs were cut short because the runner canceled the second arm whenever the first failed to finish; restoring all 27 boundary runs shows the projected method could land anywhere from 9 tasks worse to 1 task better than the full-context method. Of the fifteen completed pairs, only eleven had both arms succeed, and that smaller, fully-matched group is what the study actually used to compare resource use: projection cut total logical tokens by 25%, but the median pair still used 29% more tokens, with total follow-up requests rising from 35 to 55.
This is less a verdict on context projection than a warning about how evaluation scaffolding can hide failure: when a runner drops a companion test the moment its partner fails, real failures disappear from the scorecard instead of counting against it. The risk is concrete, not hypothetical - one dropped run had the projected-context agent burn twelve requests digging through archived text for an answer its full-context counterpart gave in three.
Context compression is marketed across agent frameworks as a free efficiency win; this study is a reminder to check what counts as "success," and how many runs got quietly excluded before the scorecard was drawn up.