Researchers have found a way to make AI agents forget the right things, without retraining anything.
AI agents that run multi-step tasks pile up long histories of tool calls, API responses, and dialogue, and that bloat both slows them down and confuses them, a known problem called attention dilution. A new paper describes FOCUS, a test-time method that identifies which past interactions actually shaped an agent's later decisions and keeps only those, discarding the rest on the fly. Unlike prior compression systems, which train a policy offline on a fixed dataset before deployment, FOCUS needs no training data or fine-tuning, and it can attach to any closed-API model as an add-on layer. Tested across tool-calling, question-answering, web-browsing, and multi-turn dialogue benchmarks, it cut peak context size by up to 48 percent and reduced dependency chains by 73 percent, while raising task success by as much as 8.9 percentage points over running the agent uncompressed.
Most context-compression work treats this as a data problem: collect examples, train a summarizer or policy, then hope it generalizes to whatever a live agent encounters later. FOCUS instead asks a causal question at run time, which steps actually caused the next decision, and drops the rest, adapting to each trajectory instead of a fixed rulebook. That is the more interesting shift: it turns compression into something the model reasons through as it works, not a preprocessing step bolted on before deployment.
The efficiency numbers are eye-catching, but they come from one paper's own benchmarks. Independent replication, and a hard look at whether "causal" attribution holds up on messier real-world agent logs, will show whether this survives outside a lab.