AI/ ai agents · reinforcement learning · coding agents · benchmarks

Small AI Coding Agents Match Models 13.6 Times Their Size

A new training method teaches small coding agents to use files as memory, letting a 9B model match one 13.6 times its size on coding benchmarks.

A 9-billion-parameter AI coding agent is now holding its own against a model 13.6 times its size, thanks to a new way of training it to take notes.

Researchers built a training suite called Coding Agent Memory Gym (CAMG), covering four long-horizon task types: shopping, coding, deep research, and autonomous research. Instead of giving the agent a custom memory tool, they gave it shell access and a persistent workspace, letting it create, edit, and search its own files as memory. They then trained a single policy across all four environments with reinforcement learning, using a method called asynchronous PPO, rewarding the agent only for finishing tasks - not for managing memory any particular way. The resulting models, CAMG-RL-4B and CAMG-RL-9B, were built from matching-size Qwen3.5 base models.

On two established benchmarks, SWE-bench Verified and MLE-bench Lite, the 9B version performed about as well as Qwen3.5-122B-A10B, a model 13.6 times larger. The 4B version matched Qwen3.5-35B-A3B, a model nearly nine times its size. That gap matters because smaller models are cheaper to run and easier to deploy without a server farm backing them up.

Benchmark parity is not real-world parity, and every claim that a small model beats a big one eventually meets a less forgiving task outside the test set.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →