[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-test-ai-training-method-with-no-reward-signal":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},9223,"researchers-test-ai-training-method-with-no-reward-signal","Researchers Test AI Training Method With No Reward Signal","A proof-of-concept architecture that rewards only survival makes reward hacking structurally unstable, the authors argue, though it remains unproven at scale.","A new self-training setup for AI systems skips reward functions altogether, letting results survive or die based on whether they actually work in the world.\n\nThe approach, described in a replaced arXiv preprint, builds a proof-of-concept architecture where an AI's candidate behaviors run under real resource constraints. There's no score, no objective function, no task-specific supervision. The only test is whether a behavior's effects on the environment persist and leave room for more interaction later. Behaviors that pass that bar stick around; everything else gets pruned, in a process the researchers call negative-space learning.\n\nReward hacking is the chronic failure mode of self-training: models learn to satisfy whatever proxy judges them, rather than the task itself. By removing the proxy and replacing it with raw survival in an environment, the authors argue reward hacking becomes structurally unstable rather than merely discouraged. Along the way, the system reportedly developed its own tricks, like deliberately failing an experiment to generate an informative error message, without being told to try that.\n\nThat's a genuinely interesting result for a proof-of-concept, and the line between \"evolutionarily unstable\" and \"doesn't happen\" is exactly where more testing needs to go before anyone builds a production system on it.","[\"ai\",\"ai-safety\",\"self-training\",\"reinforcement-learning\"]","2026-10-01T04:00:00.000Z","2026-10-02T02:43:54.075Z","2026-10-02T02:43:57.985Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The dek claims the system sidesteps reward hacking 'entirely,' but the body only reports the authors' own hedged argument that it becomes 'structurally unstable' in an unproven proof-of-concept — soften the dek to match that hedging.","resolved","ai",[30,32,33,34],"ai-safety","self-training","reinforcement-learning",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2601.12310",0,{"sections":41},[42,45,49,53,58,63,67,72,77,81,86,91,96,101],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5629,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",816,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Policy","policy",430,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":18},"Science","science",163,{"name":68,"slug":69,"count":70,"latest_published_at":71},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":78,"slug":79,"count":75,"latest_published_at":80},"Software","software","2026-09-30T21:41:11.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]