[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-use-rubrics-instead-of-rewards-to-train-ai-agents":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},9496,"researchers-use-rubrics-instead-of-rewards-to-train-ai-agents","Researchers Use Rubrics Instead of Rewards to Train AI Agents","A new training method called RISED uses text-based rubrics, not just pass-fail scores, to decide what an AI agent practices and how it learns from mistakes.","A new paper tackles a specific flaw in training all-purpose AI agents: a single pass-fail score doesn't tell you much when every rollout in a batch gets the same score.\n\nResearchers describe RISED (RubrIcs for agentic multi-environment Selection and sElf-Distillation), a method for training one LLM agent across many different interactive environments at once. Existing approaches pick training data environment-by-environment or by local reward, without looking at how rollouts across environments relate to each other. Because environments get learned at different speeds, a single batch can end up with groups where every attempt fails or every attempt succeeds - leaving no reward contrast to learn from. RISED has an LLM judge tag each rollout with a shared vocabulary of behavioral rubrics, then uses those tags to pick training data that matches the overall behavior mix in the batch while skipping data too similar to what's already selected. Positive rubrics, which describe good behavior, feed an on-policy self-distillation teacher for extra token-level supervision; negative rubrics, which describe bad behavior, steer future rollouts away from repeat failure patterns.\n\nThe pitch is that text describing what happened carries more signal than a number saying whether it happened, especially when the numbers are identical. Across multiple model backbones, RISED posted the highest mean pass rate and ranked first or second on every individual environment tested, a sign the gains aren't just one environment skewing the average.\n\nRubric-based training is a reasonable next step for multi-environment agent RL, but the paper tests a fixed, shared rubric vocabulary. Whether that vocabulary holds up once agents hit environments nobody wrote rubrics for is the question the abstract doesn't answer.","[\"ai agents\",\"reinforcement learning\",\"llm training\",\"arxiv\"]","2026-10-02T04:00:00.000Z","2026-10-02T21:57:36.597Z","2026-10-02T21:57:42.979Z","published",null,[],"ai",[26,27,28,29],"ai agents","reinforcement learning","llm training","arxiv",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.00979",0,{"sections":36},[37,40,44,48,53,58,62,67,72,77,82,87,92,97],{"name":38,"slug":24,"count":39,"latest_published_at":18},"AI",5859,{"name":41,"slug":42,"count":43,"latest_published_at":18},"Security","security",833,{"name":45,"slug":46,"count":47,"latest_published_at":18},"Policy","policy",438,{"name":49,"slug":50,"count":51,"latest_published_at":52},"Deals","deals",317,"2026-10-01T22:00:00.000Z",{"name":54,"slug":55,"count":56,"latest_published_at":57},"Hardware","hardware",198,"2026-10-01T17:38:48.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":18},"Science","science",171,{"name":63,"slug":64,"count":65,"latest_published_at":66},"Consumer Tech","consumer-tech",155,"2026-10-01T19:54:10.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Dev Tools","dev-tools",96,"2026-10-01T16:57:03.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Software","software",93,"2026-09-30T21:41:11.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":81},"Startups","startups",90,"2026-10-01T21:55:22.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]