What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks
Summary
Benchmarks are a primary way to assess and communicate progress in large language models, but rankings alone do not show how evaluation expectations are changing. This study maps 14,767 papers that introduced or updated evaluation resources in arXiv submissions published between January 2022 and August 2026. Using staged screening and automated full-text coding, the authors analyze the systems and domains being tested, the materials and conditions used for evaluation, and the scoring mechanisms applied. The analysis finds increasing emphasis on action, interaction, and professional applications, while older and newer benchmark design elements often coexist. The use of LLMs as evaluators is growing in both agent and non-agent benchmark groups. By contrast, the study finds no similarly sustained increase in model-generated evaluation materials in recent cohorts. The findings describe how capability expectations are converted into concrete tests and success criteria. They also raise a methodological concern: as AI systems help construct tests, perform tasks, and judge answers, expanding evaluation may provide less independent evidence if it reproduces the preferences or blind spots of the models involved.