AI agents trained with reinforcement learning tend to settle on one way of solving a problem, even when several equally good approaches exist.
A new paper from researchers posting on arXiv proposes a way to change that. Instead of relying on random sampling or generic regularization to produce varied behavior, the method lets developers define exactly which kind of variation matters for a task, like differing tool-use patterns or reasoning paths. The authors call this approach Trajectory-guided Joint Policy Optimization (TJPO). It works by scoring groups of sampled trajectories against user-specified descriptors, then optimizing a single policy to spread out across those descriptors rather than training multiple separate policies. Tests on the Sokoban puzzle game and the ALFWorld household-task simulator showed the method produced genuinely distinct successful strategies while keeping task performance competitive with standard training.
That distinction matters because brittle agents are a recurring problem in deployed systems: an agent that only knows one path to a goal can fail outright when something about the environment shifts, like a blocked tool or a changed interface. Explicit, controllable diversity gives an agent a menu of backup strategies instead of a single brittle script, and it gives developers a lever to shape that menu toward the kind of variation they actually want.
The catch is that Sokoban and ALFWorld are tidy, bounded benchmarks, nothing like the messy multi-step coding or browsing tasks agentic AI products are being sold on, so whether this scales past toy environments is still an open question.