A new academic survey takes stock of how optimal transport math is creeping into reinforcement learning.
The paper, posted to arXiv as paper 2610.01413 on October 2, 2026, reviews how optimal transport, or OT, is used to compare probability distributions inside reinforcement learning systems. RL algorithms constantly compare distributions - how an agent moves through states versus how an expert does, or what actions a learned policy picks versus what's in an offline dataset. The authors note that standard difference measures break down when those distributions barely overlap, a common problem in imitation learning, offline RL, and real-world deployment. OT sidesteps that by calculating the cost of shifting probability mass from one distribution to the other, using a cost function tied to the task's own geometry.
Distribution mismatch is a quiet, persistent failure mode in RL - it's a big reason policies trained in simulation flounder once they hit messier real-world data. Rather than proposing a new algorithm, the survey maps dozens of existing OT-based methods by the role OT plays, which distributions get compared, and how each handles time, turning scattered papers into something closer to a field guide.
Surveys rarely make news, but a field quietly borrowing math from transportation theory to patch a fundamental RL weakness is worth watching - especially since the authors admit that scaling OT to full trajectories remains unsolved.