Back to News
RSS feedarxiv.org

GAP-DPO Aligns Preference-Pair Selection with User Utility for Personalized LLMs

Summary

Personalizing large language models requires optimizing for individual users’ preferences rather than aggregate response quality. This paper studies how the selection of preference pairs affects Direct Preference Optimization (DPO) in that setting. The authors argue that common heuristic selection rules, including likelihood-based extremes, can become disconnected from explicit user utility and may weaken personalization. They formalize the problem geometrically by analyzing the first-order interaction between gradients of expected user utility and DPO update directions. Under off-policy sampling, their analysis shows that DPO can shift from an error-corrective update to a reinforcement-like update when preference margins align directionally with utility gradients. Based on this result, they propose GAP-DPO, an iterative method that selects pairs according to user utility and gradient geometry while limiting distribution shift through epoch-wise data regeneration. Experiments on personalized text-generation benchmarks report consistent gains over standard DPO variants in stylistic fidelity, preference alignment, and overall generation quality. The paper presents gradient alignment as a unifying principle and treats pair selection as part of the optimization geometry rather than a preprocessing heuristic.