AI/ ai · ai-safety · self-training · reinforcement-learning

Researchers Test AI Training Method With No Reward Signal

A proof-of-concept architecture that rewards only survival makes reward hacking structurally unstable, the authors argue, though it remains unproven at scale.

A new self-training setup for AI systems skips reward functions altogether, letting results survive or die based on whether they actually work in the world.

The approach, described in a replaced arXiv preprint, builds a proof-of-concept architecture where an AI's candidate behaviors run under real resource constraints. There's no score, no objective function, no task-specific supervision. The only test is whether a behavior's effects on the environment persist and leave room for more interaction later. Behaviors that pass that bar stick around; everything else gets pruned, in a process the researchers call negative-space learning.

Reward hacking is the chronic failure mode of self-training: models learn to satisfy whatever proxy judges them, rather than the task itself. By removing the proxy and replacing it with raw survival in an environment, the authors argue reward hacking becomes structurally unstable rather than merely discouraged. Along the way, the system reportedly developed its own tricks, like deliberately failing an experiment to generate an informative error message, without being told to try that.

That's a genuinely interesting result for a proof-of-concept, and the line between "evolutionarily unstable" and "doesn't happen" is exactly where more testing needs to go before anyone builds a production system on it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →