Back to News
RSS feedarxiv.org

Rasch Theory Reveals Hidden Biases in LLM Evaluation

Summary

This paper frames LLM evaluation as a measurement problem and applies many-facet Rasch models to annotations from nine LLMs. The study finds systematic differences between model and human raters in severity, calibration, question-order robustness, target-identity sensitivity, and rating-scale use. The authors recommend Rasch measurement theory for evaluating LLMs as examinees, judges, and raters.