AI/ ai-research · self-improving-ai · machine-learning · ai-safety

Researchers Propose Letting AI Train Itself With No Human Data

A new arXiv paper proposes training AI on self-generated, self-validated data using metrics like disk space or followers instead of human feedback.

A new paper wants AI models to stop studying our homework and start grading their own.

Researchers describe AI agents that generate strategies and executable code aimed at maximizing a single numeric reward, such as claimed disk space or a follower count, rather than a human-labeled benchmark. Successful runs feed back into retraining using a fine-tuning method called GRPO, while separate modules handle environment analysis, strategy generation, and code synthesis. The authors say prioritizing real-world validation over textual similarity to existing data helps dodge two known failure modes: model collapse and the "warm start problem," where a system has too little data to begin learning. The paper is a revised version of a submission first circulated in 2025.

This matters because human-generated training text is a finite resource, and labs already blend in synthetic data, risking feedback loops that degrade model quality across generations. Letting a model mine its own reward from direct interaction with an environment is a plausible way around both constraints - assuming the reward really can't be gamed.

That "assuming" is the whole pitch: handing a model an open-ended metric and the freedom to write its own code to chase it is exactly the setup security researchers spend careers trying to lock down.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →