Tropical Reinforcement Learning Improves Compositional Reasoning in Language Models
Summary
Reinforcement learning for large language models usually maximizes expected return by summing the probabilities of successful trajectories. The paper argues that this objective records how often a policy succeeds but does not preserve which verified solution worked, and reinforcing one solution can reduce the probability of another without evidence that it is incorrect. This is problematic for compositional reasoning, where useful reasoning steps may appear across separate, mostly failed rollouts rather than in one complete attempt. The authors propose Tropical Reinforcement Learning, which replaces addition over alternative solutions with a maximum operation from the tropical semiring. Each state is assigned the log-probability of its most likely verified solution together with an explicit path that can be replayed and reused. This lets the method join the best verified prefix and suffix that meet at a shared state, even when they originated in different rollouts. The resulting TROPIC algorithm is designed for deterministic, resettable environments with verifiable outcomes. On Sokoban, Countdown, FrozenLake, and WebShop, it outperforms the strongest on-policy baselines by as much as 16 percentage points. The results suggest that changing the algebra of the reinforcement-learning objective can improve compositional reasoning beyond changes to estimators alone.