AI agents that train on their own best-looking mistakes get measurably better at finishing complex tasks.
Researchers built a framework where two AI systems train together. A "target" agent tries to complete multi-step jobs like shopping online, answering science questions, and querying databases, while a second "failure agent" studies both systems' flops and learns to generate the sneakiest kind of wrong answer - ones that look almost right. The target agent then practices telling those near-misses apart from genuine successes, a technique called preference training. As the target agent improves, its new mistakes feed back into the failure agent, which gets better at manufacturing tougher near-misses, and the cycle repeats. Across three task types - online shopping, scientific reasoning, and interactive SQL querying - the setup lifted average task reward by 5.7 points over a standard baseline.
Most self-improving agents just collect whatever failures they stumble into and use them as training material, but obvious blunders teach an agent little. It is the plausible-looking errors that force it to sharpen its judgment. Manufacturing those harder cases on purpose, instead of waiting to collect them by accident, could let agents improve faster without more human-labeled data or bigger models.
The 5.7-point gain comes from the researchers' own benchmarks, and the code is not public yet, so outside labs have no way to kick the tires until it ships.