Back to News
RSS feedarxiv.org

EvalDetectBench Measures Evaluation Awareness in Frontier LLMs

Summary

Researchers introduce EvalDetectBench, an open pipeline and benchmark for measuring whether frontier language models recognize when they are being evaluated. It works with Inspect-compatible evaluations and includes transcripts from system-card assessments and deployment sources. The study identifies two sources of systematic bias: deployment-transcript generator identity explains 11.25% of measurement variance and can reorder model rankings, while prompts optimized for one model may fail on others. Per-model probe calibration and stratified generator harmonization are designed to address these problems and improve the reliability of safety evaluations.

EvalDetectBench Measures LLM Evaluation Awareness | Benpay.ai Board