Researchers found a cheap trick that makes reinforcement learning agents smarter: have them describe themselves in plain text as they go.
STRAT adds one extra prediction head to a standard reinforcement learning policy. That head learns to output a short text trace of the agent's own state - its position, inventory, goals, and immediate progress - borrowing ideas from how humans describe navigating a space using landmarks, routes, and an overall sense of layout. The traces are generated automatically by the environment's own rules, so no human labeling is required. Tested across 60 sparse-reward tasks in the XLand-MiniGrid benchmark, agents with STRAT solved environments that standard reinforcement learning could not touch at all, while also keeping their internal state representations more compact and avoiding a failure mode called rank collapse.
This lands at a moment when interpretability is reinforcement learning's weak spot: trained policies are notoriously hard to audit, and chain-of-thought explanations in language models have already shown they can be unfaithful to what the model is actually doing underneath. STRAT's trace is not a bolted-on explanation after the fact - it is trained into the objective itself, and the paper reports it improves task performance rather than just legibility. That combination, if it holds up beyond sparse-reward toy benchmarks like XLand-MiniGrid, would be rarer than it sounds, since most interpretability work trades away performance to get insight instead of getting both.
Whether a line like agent picked up key, heading to door scales to anything messier than a grid world is the open question this paper leaves for someone else.