[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-a-video-ai-reward-model-that-grades-with-rubrics-not-vibes":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},7332,"a-video-ai-reward-model-that-grades-with-rubrics-not-vibes","A Video AI Reward Model That Grades With Rubrics, Not Vibes","RewardVerse scores AI video against explicit criteria instead of one opaque number, aiming to fix the unstable reward signals that plague video-generation RL.","A new paper proposes grading AI-generated video with written rubrics instead of a single gut-check score.\n\nResearchers behind RewardVerse built a reward model for video generation that first drafts explicit evaluation criteria for a given prompt, then scores the video against those criteria rather than guessing a number outright. The training process, called Rubric-Guided Policy Optimization, runs in two stages: it warms up the scorer using self-generated seed rubrics, then jointly trains a rubric generator and scorer so the criteria adapt to each query while staying aligned with human ratings. The team tested the system on EvalVerse, a benchmark spanning 16 quality dimensions, plus outside datasets. They report the rubric-based approach curbs what they call scalar drift - the tendency for a reward model's scoring scale to shift or collapse across different prompts - and claim state-of-the-art results on both single-video and head-to-head video comparisons.\n\nReward models are the referee for reinforcement learning in video generation: get the scoring wrong and training drifts toward whatever the model can game, not what actually looks good. A rubric that adapts per query is a sensible answer to a problem RLHF researchers have wrestled with for years in text models, where unstable or exploitable reward signals produce over-optimized, hollow outputs.\n\nThe paper does not name the specific prior reward models it claims to beat or publish the score margins on EvalVerse, so \"state-of-the-art\" is a claim to take on faith until the numbers or code show up.","[\"ai\",\"video-generation\",\"reward-models\",\"reinforcement-learning\"]","2026-09-23T04:00:00.000Z","2026-09-23T07:57:58.856Z","2026-09-23T07:58:05.629Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"publisher-r1","publisher",1,"The body describes RewardVerse's two-stage 'Rubric-Guided Policy Optimization' training method but never names or benchmarks against the 'prior methods' or specifies numeric results on the 16-dimension EvalVerse benchmark, leaving the core efficacy claim unverifiable and reading as vague rather than fact-checked.","resolved","ai",[30,32,33,34],"video-generation","reward-models","reinforcement-learning",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.22947",0,{"sections":41},[42,45,49,54,59,63,67,72,77,82,87,92,97,102],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",4297,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",710,{"name":50,"slug":51,"count":52,"latest_published_at":53},"Policy","policy",369,"2026-09-23T02:13:52.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",202,"2026-09-22T23:00:04.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":18},"Hardware","hardware",169,{"name":64,"slug":65,"count":66,"latest_published_at":18},"Science","science",133,{"name":68,"slug":69,"count":70,"latest_published_at":71},"Consumer Tech","consumer-tech",110,"2026-09-22T20:00:00.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Software","software",80,"2026-09-22T23:32:52.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":81},"Dev Tools","dev-tools",79,"2026-09-22T22:21:13.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",65,"2026-09-22T22:06:48.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Gaming","gaming",45,"2026-09-22T15:35:06.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"General","general",43,"2026-09-21T23:48:56.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",27,"2026-09-22T13:00:00.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]