An AI agent can learn to rewrite another AI agent's robot-control instructions, but only if the testing setup is rigorous.
A coding agent was used as a robot's brain: it watches a browser-based 3D scene through screenshots and moves things by placing a virtual gripper with a few tools. A second "optimizer" agent rewrote that first agent's harness, its prompts, tools, and control rules, between rounds of testing. With only 5 rollouts per round, the optimizer could get lucky: held-out success started near 47 percent while training scores (70 percent) badly overstated real performance. Scaling to 100 rollouts per round lifted held-out success to 67 percent, and separately, requiring each revision to beat the prior version head-to-head, rather than accepting every edit, turned a harness that was getting worse after ten rounds into one that climbed from 51 percent to 67 percent over 30 rounds.
This matters because the harness, not the underlying model, has quietly been the bottleneck in using general-purpose coding agents as robot controllers, and until now it was hand-tuned by trial and error. The result borrows a basic lesson from machine learning, you need a large enough test set and a strict promotion gate, and applies it to agents editing other agents' instructions, suggesting self-improving agent loops need the same discipline as any training pipeline, not less.
Strip away the bigger batch size and the champion-challenger check, and the same setup degrades within ten rounds, a useful reminder that "self-improving AI" usually just means somebody finally added a proper evaluation set.