Auditing LLM Judges for Occupational AI Measurement
Summary
LLM judges are increasingly used to assess whether AI outputs satisfy workplace requirements, but this study finds that ranking agreement is not enough for occupational measurement. The authors introduce O*NET-BENCH, an audit suite derived from survey data containing 45,796 worker ratings, and evaluate 33 pre-existing judge configurations across six model families on 4,501 test ratings. Twenty-five configurations achieve tie-aware pair accuracy of at least 0.60, yet a train-fitted response-only TF-IDF baseline nearly matches the strongest judge. The configurations estimate that 3.0% to 97.9% of responses are acceptable, while occupation-matched workers report an acceptance rate of 61.1%. In one fine-tuned lineage, switching from pointwise scoring to a bundled few-shot/listwise protocol improves response ordering but reduces agreement with worker means at both task and occupation levels; the reversal also appears on a task- and worker-disjoint validation split using prespecified criteria. Cross-validated calibration largely removes mean bias, but calibrated scores explain no more than 8.5% of individual worker-rating variance. Prediction-assisted estimation provides only small precision gains at the tested label budgets. The authors conclude that judges must be validated against the acceptance rates and aggregates their scores are intended to estimate, rather than being assessed only by ranking accuracy.