[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-rl-technique-lets-ai-models-train-without-human-labels":10,"sections":39},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":34,"feedback":38,"feedback_at":22,"cost_usd":38,"total_tokens":38},8605,"new-rl-technique-lets-ai-models-train-without-human-labels","New RL Technique Lets AI Models Train Without Human Labels","An unreviewed arXiv preprint proposes a self-rewarding RL method that trains AI models without human labels, but peer review is still pending.","A new preprint claims AI models can train themselves without any human-labeled data, by learning to grade their own answers more reliably.\n\nThe paper, posted to arXiv as arXiv:2609.36750, describes a technique called Group-Marginalized Advantage Estimation, or GMAE. It's a self-rewarding reinforcement learning method: instead of humans scoring an AI's responses, the model scores itself by comparing outputs within randomly sampled groups of its own answers. The authors argue that comparing a response to just one group of peers is a noisy, incomplete way to judge quality, so GMAE averages reward signals across many possible group contexts instead. The preprint reports testing this across eight benchmarks and four base models, claiming stronger results and low added computational cost compared to existing ensemble-based self-rewarding methods.\n\nSelf-rewarding RL matters because human labeling is the single biggest bottleneck and cost driver in training capable AI models. If a method like GMAE genuinely reduces noise in self-generated reward signals, it could make self-improving training loops more stable and cheaper to run at scale. That's a meaningful technical claim, not just an incremental tweak, though it's also exactly the kind of result that needs independent replication before anyone bets a training run on it.\n\nThe paper lists no named authors or institutional affiliation and has not been peer-reviewed, so treat these benchmark numbers as a promising claim rather than a proven one.","[\"ai\",\"reinforcement-learning\",\"llm-training\"]","2026-09-30T04:00:00.000Z","2026-09-30T14:22:53.181Z","2026-09-30T14:22:59.124Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Attribute the findings explicitly to the arXiv preprint (cite arXiv:2609.36750) and note it's an unreviewed preprint with no named authors\u002Finstitution given, since the body currently states the benchmark and cost claims without naming the source publication.","resolved","ai",[30,32,33],"reinforcement-learning","llm-training",[35],{"name":36,"url":37},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.36750",0,{"sections":40},[41,44,48,52,57,62,67,72,77,81,86,91,96,101],{"name":42,"slug":30,"count":43,"latest_published_at":18},"AI",5135,{"name":45,"slug":46,"count":47,"latest_published_at":18},"Security","security",788,{"name":49,"slug":50,"count":51,"latest_published_at":18},"Policy","policy",417,{"name":53,"slug":54,"count":55,"latest_published_at":56},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":58,"slug":59,"count":60,"latest_published_at":61},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":63,"slug":64,"count":65,"latest_published_at":66},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":82,"slug":83,"count":84,"latest_published_at":85},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]