AI/ ai agents · self-improving ai · llm benchmarks · test-time compute

Researchers Build AI Agents That Self-Improve Over Many Rounds

A new research system trains AI agents to repeatedly critique and rewrite their own answers, and the benchmark gains keep climbing instead of leveling off.

A new agent framework called AREX-2 teaches AI models to grade and rewrite their own answers across dozens of rounds, and the gains keep compounding instead of plateauing.

Researchers trained the agent on synthetic "long-horizon improvement trajectories" drawn from machine learning and algorithmic programming tasks, domains where right and wrong answers are easy to verify. The bet is that two skills learned there - spotting a better solution (reflection) and sticking with the grind over many iterations (long-horizon execution) - should transfer to other kinds of work. Built on Alibaba's open-weight Qwen3 model family, the resulting agent scored 81.8 on MLE-bench Lite and 70.7 on Frontier-CS, two machine-learning and coding benchmarks. It also carried over to tasks it wasn't explicitly trained on: 84.0 on BrowseComp, 52.6 on HLE (Humanity's Last Exam, a notoriously difficult general-knowledge test), 92.2 on GAIA (a benchmark of real-world assistant tasks), and 93.8 on DeepSearchQA.

The interesting part isn't the raw scores - it's that performance kept climbing as the agent got more rounds to iterate, rather than hitting a wall after a couple of tries, which is the usual failure point for self-correcting agents. Teaching reflection in verifiable domains like code and ML, then watching it transfer to messier, open-ended research tasks, backs a broader industry wager: that letting models think longer at test time, not just training bigger ones, is where the next gains come from.

It's one paper's benchmarks, not an independent audit, and self-improving agents have a habit of acing curated test sets while stumbling on messier real-world workloads - so treat the transfer claims as promising, not proven.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →