The paper introduces Group Variance Policy Optimization (GVPO), a post-training method for large language models designed to address instability caused by importance sampling in existing approaches such as Group Relative Policy Optimization. GVPO incorporates the analytical solution to KL-constrained reward maximization into its gradient weighting scheme. The authors interpret its gradient as minimizing the squared difference between the centered implicit and actual rewards. They state that GVPO has a unique optimum matching the KL-constrained reward-maximization objective. It also permits flexible sampling distributions without importance sampling. The method naturally extends to on-policy distillation and can optimize a broad family of extended on-policy distillation objectives. The paper presents these properties as a foundation for more reliable and adaptable LLM post-training and distillation.
AI News
The latest AI releases, research, products, and industry updates.
Loading...