NashEval Learns Context-Specific Equilibria for Robust Agent Evaluation
Summary
The paper studies how to evaluate AI agents when judgments depend on context, such as a prompt, task, or user group, and when human preferences are heterogeneous. Conventional score-based approaches such as Bradley-Terry impose a transitive ordering, which may not represent collective preferences with cycles or disagreement. The authors instead formulate evaluation as a contextual two-player game in which each player chooses a distribution over agents, and define the support of the Nash equilibrium as the context-specific winner set. Learning these equilibria from offline logs is difficult because each context usually contains feedback on only a subset of agents, making a naive plug-in payoff estimate biased. NashEval addresses this by first constructing debiased estimates of the contextual payoff matrix. It then learns the mapping from context to equilibrium with a tailored orthogonal loss, avoiding the need to solve a separate game for every context. Theoretical analysis shows that errors in the nuisance functions used to estimate the payoff matrix affect the learned equilibrium's exploitability only through higher-order terms. Across multiple experiments, NashEval improves the robustness of equilibrium learning and consistently identifies the top-performing agents for different contexts.