AI/ ai · ai agents · benchmarks · model training

AI Agent RSI-Master Beats Human-Tuned Qwen3 Model on Coding Test

RSI-Master, an AI agent that runs its own training experiments, beat Qwen3's human-tuned 35B Instruct model on two benchmarks.

An AI agent that designs its own training experiments just outperformed a model humans fine-tuned by hand.

RSI-Master is built to let an agent run open-ended post-training experiments on a base language model without a human steering every step - and without quietly cutting corners to inflate its own scores. To stop that kind of gaming, the researchers added an "Experiment OS" that logs every action in a traceable record, plus a reviewer system that compares results across a growing tree of experiments so the agent doesn't lock onto one approach too early. On PostTrainBench, using a 4-billion-parameter Qwen3 base model, RSI-Master averaged a score of 54.49 against 46.53 for the best rival agent, with a reported 0.0% hacking rate. Scaled up to a 35-billion-parameter model, RSI-Master's resulting model beat Qwen3's own human-tuned Instruct version on the LiveCodeBench-v6 coding benchmark (41.21 vs. 37.36) and on SciCode, and posted a nonzero score on HorizonMath, a test of unsolved research problems where most frontier models score close to zero.

The real finding isn't the benchmark scores - it's the zero hacking rate. Letting an AI system redesign its own training process is only useful if it isn't also learning to cheat the metrics that measure success, and that's the failure mode this architecture is explicitly built to catch. If the approach holds up at larger scale, it suggests autonomous post-training can be made auditable, not just fast.

A nonzero score on HorizonMath sounds unimpressive until you remember most frontier models can't clear that bar at all.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →