[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-jev-a-frozen-decision-model-now-helps-train-rl-agents":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},11020,"jev-a-frozen-decision-model-now-helps-train-rl-agents","Jev, a Frozen Decision Model, Now Helps Train RL Agents","A new study shows the decision model Jev, used without any extra training, improves reinforcement learning across nine MiniGrid tasks and three Atari games.","A frozen AI model just made reinforcement learning work better without being trained itself.\n\nResearchers tested whether Jev, a decision model that returns calibrated answers in one forward pass instead of generating text token by token, could speed up reinforcement learning. They found it could fill nearly every role an RL system needs, apart from acting as a value function, and built three ways to plug it in: as a reference policy, an exploration judge, and a replay rater. Across nine MiniGrid tasks and three Atari games, adding Jev to a standard RL learner improved results, including cases where the learner made no progress on its own. Jev itself stayed frozen throughout, never updated or fine-tuned for the job.\n\nReinforcement learning has long struggled to learn efficiently from scratch, and foundation models are the obvious fix but are usually too slow and expensive to query token by token during training. This work shows a model built to answer instantly, not generate, can plug straight into existing RL loops as a judge or rater rather than a thing to fine-tune. That is a cheaper, more practical path to borrowing foundation model knowledge than retraining or prompting a chatbot at every step.\n\nIt is a frozen, untrained component doing meaningful work, which says as much about how limited standard RL still is as it does about Jev.","[\"reinforcement-learning\",\"foundation-models\",\"ai-research\",\"jev\"]","2026-10-09T04:00:00.000Z","2026-10-10T02:24:11.431Z","2026-10-10T02:24:15.839Z","published",null,[],"ai",[26,27,28,29],"reinforcement-learning","foundation-models","ai-research","jev",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.11692",0,{"sections":36},[37,40,44,49,54,58,62,67,72,77,81,86,91,96],{"name":38,"slug":24,"count":39,"latest_published_at":18},"AI",6733,{"name":41,"slug":42,"count":43,"latest_published_at":18},"Security","security",931,{"name":45,"slug":46,"count":47,"latest_published_at":48},"Policy","policy",486,"2026-10-08T22:40:11.000Z",{"name":50,"slug":51,"count":52,"latest_published_at":53},"Deals","deals",474,"2026-10-08T22:00:00.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":18},"Hardware","hardware",232,{"name":59,"slug":60,"count":61,"latest_published_at":18},"Science","science",193,{"name":63,"slug":64,"count":65,"latest_published_at":66},"Consumer Tech","consumer-tech",181,"2026-10-08T23:26:35.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Startups","startups",117,"2026-10-08T16:45:00.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Software","software",114,"2026-10-08T17:57:01.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":18},"Dev Tools","dev-tools",106,{"name":82,"slug":83,"count":84,"latest_published_at":85},"General","general",66,"2026-10-09T04:46:11.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Gaming","gaming",58,"2026-10-08T20:08:45.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Reviews","reviews",34,"2026-10-08T14:00:22.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"How-To","how-to",8,"2026-10-05T09:00:00.000Z"]