Back to News
RSS feedarxiv.org

Learning Heterogeneous Human Preferences for AI Reward Models

Summary

Human feedback is widely used to train AI systems by learning a model of human utility and using it as a reward model. The authors argue that the common assumption of one universal utility function fails in subjective domains, where people can have different but internally consistent preferences. Drawing on rational choice theory, they introduce “individuated” utility functions conditioned on both the person making a choice and the decision context. They also propose a multi-stage architecture that estimates these functions from multimodal data. The approach was evaluated on a newly collected dataset containing more than 575,000 pairwise aesthetic judgments from 2,398 participants comparing automotive wheel designs. Individuated utility models substantially outperformed universal utility models, including foundation-model baselines, in the reported experiments. The findings suggest that annotator disagreement in subjective tasks can represent meaningful preference heterogeneity rather than simple annotation noise. The authors conclude that collecting annotator attributes and modeling whose preferences a reward model represents could help AI systems better capture human decision diversity.