A 9-billion-parameter AI coding agent is now holding its own against a model 13.6 times its size, thanks to a new way of training it to take notes.
Researchers built a training suite called Coding Agent Memory Gym (CAMG), covering four long-horizon task types: shopping, coding, deep research, and autonomous research. Instead of giving the agent a custom memory tool, they gave it shell access and a persistent workspace, letting it create, edit, and search its own files as memory. They then trained a single policy across all four environments with reinforcement learning, using a method called asynchronous PPO, rewarding the agent only for finishing tasks - not for managing memory any particular way. The resulting models, CAMG-RL-4B and CAMG-RL-9B, were built from matching-size Qwen3.5 base models.
On two established benchmarks, SWE-bench Verified and MLE-bench Lite, the 9B version performed about as well as Qwen3.5-122B-A10B, a model 13.6 times larger. The 4B version matched Qwen3.5-35B-A3B, a model nearly nine times its size. That gap matters because smaller models are cheaper to run and easier to deploy without a server farm backing them up.
Benchmark parity is not real-world parity, and every claim that a small model beats a big one eventually meets a less forgiving task outside the test set.