A new training trick lets AI coding agents forget most of what they've seen and still fix almost as many bugs.
The method, called LOHA (Latent Observations, Hard Actions), compresses older tool outputs, like error logs, file diffs, and test results, into compact soft tokens while keeping the agent's own reasoning and the last few observations as plain text. A companion training method, Anchored Context Distillation, teaches the agent to read that compressed history without drifting from how the original, uncompressed model behaves. On SWE-bench Verified, a benchmark built from real GitHub issues, this cut context size per call by 43% for a Qwen3-4B agent and 57% for a fine-tuned SWE-Master-4B-RL agent. Resolve rates dipped only slightly, from 14.5% to 12.1% for Qwen3 and from 27.5% to 21.8% for SWE-Master.
The advantage grows when memory is tight. Under a 32K-token limit, the compressed Qwen3 agent resolved 21.1% of a test subset versus 11.1% for the same agent running on full text, and in concurrent single-GPU serving it handled 1.9 times the request volume of its full-text counterpart. That is a direct answer to the real bottleneck in scaling coding agents: not raw intelligence, but how many can fit on the same hardware at once.
It is worth noting every compressed configuration still resolved fewer issues than its uncompressed twin, and the results so far cover only two 4B-parameter models on one benchmark, a preprint, not a shipped product.