Language models that rewrite their own memory file just beat the best hand-built context-management systems at their own game.
A new paper describes Context Language Models, or CLMs: instead of an external script deciding what to keep, summarize, or discard as a conversation grows, the model itself edits a running context file, unprompted. The researchers tested this zero-shot on existing models across several benchmarks. On BrowseComp-Plus, which scores how accurately an agent can dig up specific facts through multi-step web research, CLMs scored 11.4% higher while using 21.5% fewer FLOPs. On EdgeBench, a 12-hour test of an agent's ability to stay on task over a long session, CLMs gained 5% accuracy on 59% fewer FLOPs. On a 24-hour multi-repository coding task run by a swarm of agents, the improvement over standard compute budgets was 65% larger. The team also found the approach can be steered with plain-language instructions, and a dedicated reinforcement-learning setup pushed one model's BrowseComp-Plus score up 47.6% while using 12% less compute.
The real shift here is architectural, not just a score bump. Context management (deciding what an AI remembers, forgets, or hands off to another agent) has mostly lived in the surrounding harness: retrieval scripts, summarizers, prompt engineers. Moving that judgment inside the model itself means it can be trained and improved like any other skill, and it scales naturally to multi-agent systems where several context files need to coexist and stay in sync.
That's a meaningful efficiency story for anyone running long, expensive agent sessions, since fewer FLOPs per unit of accuracy is real money at scale. But it's one paper's benchmarks, on its own chosen tasks, with no independent reproduction yet. Whether letting a model manage its own memory holds up outside a research benchmark, or just moves the failure modes somewhere harder to audit, is the next question.