[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-method-makes-rl-training-for-ai-image-models-far-faster":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},6962,"new-method-makes-rl-training-for-ai-image-models-far-faster","New Method Makes RL Training for AI Image Models Far Faster","A new study finds that how you estimate likelihood, not the loss function, is what actually makes reinforcement learning work for diffusion image models.","A new paper argues the real fix for training AI image generators with reinforcement learning isn't a fancier loss function, it's estimating likelihood correctly.\n\nReinforcement learning has become popular for fine-tuning diffusion and flow-based image models, but these models don't have a computable likelihood the way language models do, so RL methods borrowed from LLM training need a workaround. Most prior research focused on inventing new policy-gradient objectives while treating the likelihood estimator as an afterthought, without testing whether the estimator itself was limiting performance. Researchers tested Stable Diffusion 3.5 Medium and separated the problem into three independent parts: the policy-gradient objective, the likelihood estimator, and the rollout sampling scheme. Swapping in an evidence lower bound (ELBO) estimator based only on the final generated image, rather than tweaking the loss formula, turned out to be what actually made training effective, efficient, and stable.\n\nThat's a meaningful reversal for a subfield that has mostly competed on inventing cleverer loss functions. If the estimator was the real lever all along, some of the gains claimed by prior fine-tuning papers may have more to do with incidental estimator choices than their headline objectives. The efficiency gap is large: the method pushed the GenEval score from 0.24 to 0.95 in 90 GPU hours, 4.6 times faster than FlowGRPO and twice as fast as the best prior method that avoids reward hacking.\n\nReward hacking still shows up as the caveat in that last comparison, a reminder that these models are often better at gaming a benchmark than at following the instruction it's supposed to measure.","[\"reinforcement-learning\",\"diffusion-models\",\"ai-research\",\"image-generation\"]","2026-09-18T04:00:00.000Z","2026-09-19T00:40:29.358Z","2026-09-19T00:40:41.276Z","published",null,[],"ai",[26,27,28,29],"reinforcement-learning","diffusion-models","ai-research","image-generation",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2602.04663",0,{"sections":36},[37,40,44,49,54,58,62,67,71,76,81,86,91,96],{"name":38,"slug":24,"count":39,"latest_published_at":18},"AI",4082,{"name":41,"slug":42,"count":43,"latest_published_at":18},"Security","security",661,{"name":45,"slug":46,"count":47,"latest_published_at":48},"Policy","policy",339,"2026-09-17T12:00:00.000Z",{"name":50,"slug":51,"count":52,"latest_published_at":53},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":18},"Hardware","hardware",155,{"name":59,"slug":60,"count":61,"latest_published_at":18},"Science","science",125,{"name":63,"slug":64,"count":65,"latest_published_at":66},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":18},"Dev Tools","dev-tools",78,{"name":72,"slug":73,"count":74,"latest_published_at":75},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]