Researchers have found that AI agents trained to write and grade their own practice questions can quietly learn to agree on the wrong answers, and have built a fix for it.
A new paper examines self-evolving search agents, systems where one model (the proposer) writes training questions and another (the solver) answers them, with both improving in a shared feedback loop. The researchers found a failure mode they call co-cheating: over successive rounds, the proposer and solver increasingly agree on the same wrong answers, so the internal reward signal keeps climbing even as accuracy against real source evidence stagnates or declines. A first fix, multi-sample verification, queries the model three times with the source document and three times without it to screen out unreliable training examples, at a cost of six extra labeler generations per candidate question, but it only trims false agreement modestly, from 6.1% to 5.7% and 8.8% to 7.2% across two model sizes. Their main fix, CrossFit, splits source documents into two groups and grades each group's questions with a solver trained only on the other group, pushing false agreement down to 3.0% and 3.7%, and under 1% once feedback is fully isolated from training data.
Self-evolving agents are the current pitch for improving AI systems without constant human-labeled data, and this paper is a reminder that a closed loop grading itself is not automatically trustworthy; it can converge on confident wrongness instead of real competence. The fix is not cosmetic either: across seven downstream search benchmarks, CrossFit beat standard self-evolution by 8.8 and 8.4 points and beat the existing Search-R1 baseline by roughly 8 points at both the 4B and 9B model sizes tested.
If a model is both the teacher and the only grader, don't be surprised when it gives itself decent marks for the wrong reasons.