AI/ reinforcement-learning · foundation-models · ai-research · jev

Jev, a Frozen Decision Model, Now Helps Train RL Agents

A new study shows the decision model Jev, used without any extra training, improves reinforcement learning across nine MiniGrid tasks and three Atari games.

A frozen AI model just made reinforcement learning work better without being trained itself.

Researchers tested whether Jev, a decision model that returns calibrated answers in one forward pass instead of generating text token by token, could speed up reinforcement learning. They found it could fill nearly every role an RL system needs, apart from acting as a value function, and built three ways to plug it in: as a reference policy, an exploration judge, and a replay rater. Across nine MiniGrid tasks and three Atari games, adding Jev to a standard RL learner improved results, including cases where the learner made no progress on its own. Jev itself stayed frozen throughout, never updated or fine-tuned for the job.

Reinforcement learning has long struggled to learn efficiently from scratch, and foundation models are the obvious fix but are usually too slow and expensive to query token by token during training. This work shows a model built to answer instantly, not generate, can plug straight into existing RL loops as a judge or rater rather than a thing to fine-tune. That is a cheaper, more practical path to borrowing foundation model knowledge than retraining or prompting a chatbot at every step.

It is a frozen, untrained component doing meaningful work, which says as much about how limited standard RL still is as it does about Jev.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →