A new paper wants AI models to stop studying our homework and start grading their own.
Researchers describe AI agents that generate strategies and executable code aimed at maximizing a single numeric reward, such as claimed disk space or a follower count, rather than a human-labeled benchmark. Successful runs feed back into retraining using a fine-tuning method called GRPO, while separate modules handle environment analysis, strategy generation, and code synthesis. The authors say prioritizing real-world validation over textual similarity to existing data helps dodge two known failure modes: model collapse and the "warm start problem," where a system has too little data to begin learning. The paper is a revised version of a submission first circulated in 2025.
This matters because human-generated training text is a finite resource, and labs already blend in synthetic data, risking feedback loops that degrade model quality across generations. Letting a model mine its own reward from direct interaction with an environment is a plausible way around both constraints - assuming the reward really can't be gamed.
That "assuming" is the whole pitch: handing a model an open-ended metric and the freedom to write its own code to chase it is exactly the setup security researchers spend careers trying to lock down.