A new preprint claims AI models can train themselves without any human-labeled data, by learning to grade their own answers more reliably.
The paper, posted to arXiv as arXiv:2609.36750, describes a technique called Group-Marginalized Advantage Estimation, or GMAE. It's a self-rewarding reinforcement learning method: instead of humans scoring an AI's responses, the model scores itself by comparing outputs within randomly sampled groups of its own answers. The authors argue that comparing a response to just one group of peers is a noisy, incomplete way to judge quality, so GMAE averages reward signals across many possible group contexts instead. The preprint reports testing this across eight benchmarks and four base models, claiming stronger results and low added computational cost compared to existing ensemble-based self-rewarding methods.
Self-rewarding RL matters because human labeling is the single biggest bottleneck and cost driver in training capable AI models. If a method like GMAE genuinely reduces noise in self-generated reward signals, it could make self-improving training loops more stable and cheaper to run at scale. That's a meaningful technical claim, not just an incremental tweak, though it's also exactly the kind of result that needs independent replication before anyone bets a training run on it.
The paper lists no named authors or institutional affiliation and has not been peer-reviewed, so treat these benchmark numbers as a promising claim rather than a proven one.