[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-training-method-teaches-ai-models-to-know-when-theyre-guessing":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8902,"new-training-method-teaches-ai-models-to-know-when-theyre-guessing","New Training Method Teaches AI Models to Know When They're Guessing","A new reinforcement learning method called RLCD helps AI models know when to trust their own answers, beating prior calibration techniques on reasoning tasks.","A new training method teaches AI reasoning models to be honest about how sure they actually are.\n\nResearchers released an open-source implementation of RLCD, short for reinforcement learning for calibrated decisions. Instead of only rewarding a model for landing on the right answer, RLCD has the model write out its reasoning, then scores the probability it ultimately assigns to that answer against what actually happened. The team found that training this way directly either shuts off reasoning entirely or gets lost in noise, so they settled on a two-step process: calibrate first, then reinforce. Tested on Qwen3-1.7B, a small open-source model, across two reasoning benchmarks, RLCD matched or beat supervised fine-tuning, STaR (Self-Taught Reasoner, a method where a model learns by generating and filtering its own rationales), and GRPO (Group Relative Policy Optimization, a reinforcement learning technique that scores answers by comparing them to a group of other sampled answers, widely used in current reasoning models).\n\nCalibration is not a side issue. Reasoning models are increasingly trained with reinforcement learning from verifiable rewards, the approach behind many recent gains in math and coding performance, but that approach has a known side effect: it makes models more accurate and more overconfident at the same time. A technique that keeps the accuracy gains while fixing the overconfidence matters for anyone using model probabilities to decide when to trust an answer versus flag it for a human.\n\nThe paper is also unusually candid about its limits: when uncertainty comes from humans disagreeing rather than the model being unsure, RLCD performs no better than plain cross-entropy training, a rare admission in a field prone to overselling.","[\"ai\",\"reinforcement-learning\",\"calibration\",\"reasoning-models\"]","2026-10-01T04:00:00.000Z","2026-10-01T10:04:21.507Z","2026-10-01T10:04:27.744Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Define the acronyms STaR and GRPO in plain language before using them as benchmark comparisons, since the draft currently name-drops both without ever explaining what they are.","resolved","ai",[30,32,33,34],"reinforcement-learning","calibration","reasoning-models",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.38850",0,{"sections":41},[42,45,50,55,60,65,70,75,80,84,89,94,99,104],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5350,{"name":46,"slug":47,"count":48,"latest_published_at":49},"Security","security",801,"2026-09-30T22:18:23.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",429,"2026-10-01T02:26:17.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Science","science",157,"2026-09-30T15:00:56.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":81,"slug":82,"count":78,"latest_published_at":83},"Software","software","2026-09-30T21:41:11.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]