AI/ reinforcement-learning · llm-reasoning · ai-research · arxiv

Tropical Reinforcement Learning Tackles AI's Compositional Gap

Researchers swap addition for maximum in RL math, letting models reuse successful reasoning steps instead of accidentally forgetting them.

A new training method changes the basic math behind reinforcement learning to help AI models handle multi-step puzzles.

Researchers propose Tropical Reinforcement Learning, which replaces the standard practice of summing probabilities across all successful solution paths with taking the single maximum probability path instead, a shift drawn from what mathematicians call tropical algebra. The team built a training algorithm called TROPIC for deterministic, resettable environments with verifiable outcomes. It lets the system splice the best first half of one attempt with the best second half of another, even when no single rollout ever produced the complete solution. Tested on four tasks (Sokoban, Countdown, FrozenLake, and WebShop), TROPIC beat the strongest on-policy baselines by up to 16 percentage points.

Standard reinforcement learning tends to reinforce one correct answer while quietly eroding a model's memory of other valid approaches, since probabilities have to sum to one. That is a particular problem for compositional reasoning, where a correct solution is built from steps a model has shown separately but rarely strings together in one pass. Tracking the single best verified path through a problem, rather than averaging across all of them, keeps useful partial solutions intact instead of letting training quietly overwrite them.

It is a narrow proof of concept across four toy-scale environments, but it suggests some of AI's reasoning limits come from the underlying math, not just model size or training data.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →