Exploration-Preserving Policy Optimization for Verifiable-Reward Reasoning
Summary
Reinforcement learning with verifiable rewards can improve reasoning, but the way learning credit is allocated affects which solutions remain discoverable through repeated sampling. The paper introduces Exploration-Preserving Policy Optimization (ExPPO), a lightweight advantage-shaping rule that redistributes credit according to prompt-relative, length-normalized response surprisal and prompt pass rate. Its bounded shaping and shared normalization are designed to preserve the sign of verifier feedback while approximately maintaining the total absolute sequence-advantage mass within each prompt group. The analysis connects response-level credit allocation with sampled-mode updates and derives local conditions under which entropy and discovery of correct modes can improve. Experiments report better in-domain and out-of-domain reasoning coverage, higher aggregate response accuracy, and strong coverage at larger sampling budgets. In a controlled multi-answer evaluation, ExPPO also increases the number of correct modes found and the diversity of verified-correct responses. The authors release code in a GitHub repository.