A new technique aims to stop AI agents from getting confused when they compress their own memory to keep long tasks running.
Researchers built PAIR (Prompt Adaptation using Interventional Rollouts), a method for agents that must periodically summarize their own interaction history to avoid overflowing their context window. Their key finding: compression doesn't make an agent suddenly fail a task outright. It first quietly lowers reliability, and the damage clusters around a handful of individual compression events rather than spreading evenly across a run. To isolate those events, the team ran matched "counterfactual continuations" - replaying the same agent from the same state with and without a specific compression step, then comparing what happened next. PAIR uses that comparison to flag which compressions hurt execution and rewrites the offending sections of a shared compression prompt template, without touching the underlying agent itself.
This matters because most existing compression-tuning methods only compare full-length runs against compressed ones after the fact, a noisy signal that gets tangled up with an agent's normal run-to-run randomness. By comparing from the same state instead, PAIR can trace an error to one specific compression decision rather than guessing across an entire trajectory - closer to debugging a crash than auditing a report card.
The catch: the abstract claims PAIR delivers the strongest cross-run reliability "in every main benchmark-scope combination" and gets compressed runs "close to" or sometimes past uncompressed baselines, but it names neither the benchmarks nor the actual success-rate numbers behind those comparisons - the kind of specifics that would let anyone outside the lab check the claim.