AI/ ai-agents · context-management · llm-costs · arxiv-research

AI Agents That Forget Tool Output Cut Costs Sharply

A new technique has AI agents archive bulky tool outputs and recall them only when needed, cutting prompt tokens by about 75 percent in one test.

AI agents that use tools - APIs, debuggers, log viewers - tend to drag every raw result along in their context window, even after they've used it once. A new preprint describes a method that lets the agent itself decide what to archive: it swaps a bulky tool result for a short note at the same spot, stashes the original in a recoverable archive, and can pull the full version back if it turns out to matter later. A lightweight Python harness handles the archiving and recovery without any retraining, and it's built to leave user instructions and the agent's own messages untouched.

The researchers tested this on an exploratory OpenTelemetry debugging session followed by an unrelated task. The forgetting agent finished with 231,951 provider-reported prompt tokens, versus 912,492 for a version that kept everything - roughly a 75 percent cut in final token count. Across the whole run, the forgetting approach also used 50 percent fewer cumulative input tokens, a separate measure of total token traffic, which brought the estimated API bill down to $1.28-$1.44 from around $4.38.

That's a real cost reduction, not a marketing one, and it matters because tool-heavy agents are exactly the kind of workload that quietly inflates cloud bills with repeated debugging logs and API dumps. But the savings came with tradeoffs: the leaner agent made more requests and took 17 percent longer in wall-clock time, and while both versions passed a basic task check, neither one satisfied the harder follow-up evaluation.

The method isn't universal. A separate pair of application-development tasks saw no token or cost savings at all, and an earlier continuation task actually scored lower on manual quality review despite using less context. So this looks less like a free lunch and more like a tool that pays off specifically when an agent is wading through noisy, repetitive logs - not when it's just writing code.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →