Training an AI to work a computer terminal is harder to pull off than it sounds, and a new paper lays out why.
Researchers built a meta-agent pipeline that uses a frontier model, Claude Opus, to generate terminal-based tasks and automated verifiers for reinforcement learning training. They found that a runnable Docker image and a passing test suite are not proof the pipeline actually works end to end. The team traced failures to three sources: invalid benchmarks, brittle test harnesses, and reward signals that do not match real task success. After redesigning prompts and extending context windows, baseline solvability of the generated tasks rose 5.6 times.
But that fix turned out to be narrow. A 9-billion-parameter model plateaued at 81.3% mean pass@2 within 20 training steps on the Opus-generated tasks. Add harder tasks to the same set, without touching the training setup at all, and mean pass@2 collapsed to 20.6%. That is strong evidence the difficulty band was calibrated to one specific model, not some universal standard.
It is a useful check on the idea that one AI can grade and author homework for another AI without a human auditing the gradebook.