Validating LLM Judges When Annotation Overlap Is Sparse
Summary
The paper studies how to validate LLM-as-a-judge systems when annotation budgets do not allow every item to receive multiple human labels. It identifies pairwise annotation overlap as the first-order factor in erroneous deployment decisions. With only 5% pairwise overlap, the reported wrong-decision rate reaches 25%, while the chance of selecting the wrong best judge from ten candidates reaches 65%. The authors derive a minimum-overlap condition of rho >= 0.25 for non-borderline judges, although borderline judges remain intrinsically difficult to distinguish. They also propose a zero-cost stratified allocation scheme that uses informative strata and halves false-rejection rates compared with random sampling. The results are validated on ten LLM judges across four evaluation matrices covering visual assessment, causal reasoning, and summarization.