AI/ reinforcement-learning · ai-research · interpretability · xland-minigrid

New RL Training Trick Has Agents Narrate Their Own State

A new training add-on has RL agents narrate their own state in plain text, and that alone solves tasks standard RL fails at outright.

Researchers found a cheap trick that makes reinforcement learning agents smarter: have them describe themselves in plain text as they go.

STRAT adds one extra prediction head to a standard reinforcement learning policy. That head learns to output a short text trace of the agent's own state - its position, inventory, goals, and immediate progress - borrowing ideas from how humans describe navigating a space using landmarks, routes, and an overall sense of layout. The traces are generated automatically by the environment's own rules, so no human labeling is required. Tested across 60 sparse-reward tasks in the XLand-MiniGrid benchmark, agents with STRAT solved environments that standard reinforcement learning could not touch at all, while also keeping their internal state representations more compact and avoiding a failure mode called rank collapse.

This lands at a moment when interpretability is reinforcement learning's weak spot: trained policies are notoriously hard to audit, and chain-of-thought explanations in language models have already shown they can be unfaithful to what the model is actually doing underneath. STRAT's trace is not a bolted-on explanation after the fact - it is trained into the objective itself, and the paper reports it improves task performance rather than just legibility. That combination, if it holds up beyond sparse-reward toy benchmarks like XLand-MiniGrid, would be rarer than it sounds, since most interpretability work trades away performance to get insight instead of getting both.

Whether a line like agent picked up key, heading to door scales to anything messier than a grid world is the open question this paper leaves for someone else.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →