AI/ ai-agents · machine-learning · recommender-systems · coding-agents

New Harness Makes AI Research Agents Check Each Other's Work

RankEvolve pits AI coding agents like Claude Code and Codex against each other to catch silent bugs that corrupt long-running ML experiments.

A new research framework makes AI coding agents double check each other's experiments, catching the kind of silent bugs that can invalidate weeks of machine learning work.

RankEvolve is a system that automates the loop of proposing, coding, training, and evaluating changes to ranking models used in applications like recommendation systems. Instead of trusting one AI coding assistant to get every step right, RankEvolve wires multiple agent products, including Claude Code and Codex, into a workflow where the agents review and repair each other's code. In a controlled, budget-matched test, that setup raised full-accuracy execution results from 45.8 percent for the best single agent to 62.5 percent, a 16.7 point gain. The researchers also ran RankEvolve for twelve iterations on HSTU (Hierarchical Sequential Transduction Units, an open-source recommender model), where it lifted NDCG@10 (Normalized Discounted Cumulative Gain at rank 10, a standard measure of ranking quality) by 4.48 percent on the MovieLens-20M LARGE benchmark and 2.80 percent on BASE.

The real story here isn't the recommendation engine gains, it's the failure mode RankEvolve targets. Long, unsupervised experiment loops can quietly leak held-out test data, skip a normalization step, or leave an evaluation flag miswired, and a single coding agent often won't catch it. Cross-checking agents cut the silent critical-defect rate to 10.4 percent, a meaningful improvement over trusting one model's word for it.

Call it mutual review for machine learning research: if one bot can slip a broken experiment past review, the bet is that two bots arguing with each other can't.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →