Rollout Efficiency in Reinforcement Learning for Reasoning Language Models
Summary
Reinforcement learning helps large language models handle mathematical, coding, and other multi-step reasoning tasks, but trajectory generation during rollout can account for a substantial share of training cost. This survey organizes recent work on making rollouts more efficient while preserving the freshness, consistency, and statistical validity of training data. It proposes a taxonomy based on both the mechanisms used and the bottlenecks they address. The authors examine how different technique families target distinct sources of inefficiency and where combining them may create opportunities or conflicts. The survey also identifies shortcomings in how efficiency gains are evaluated and reported. It concludes by outlining open challenges and future research directions for rollout systems in reasoning-oriented reinforcement learning.