UK AISI and EvalEval Open Infrastructure for Reproducible AI Evaluations
Summary
The EvalEval Coalition says the UK AI Security Institute (AISI) is using EvalEval infrastructure to openly publish evaluation results in a more reproducible and verifiable format. The collaboration builds on earlier joint research and feedback from AISI that helped shape the Every Eval Ever (EEE) reporting schema. EvalEval provides EEE and an open Evaluation Cards platform that records benchmark metadata, evaluation-run data, model metadata, and interpretive context in a common structure. AISI is sharing publicly reportable methods and findings where appropriate, including verified results, context, and configuration details for five benchmarks used in the main experiment of its paper: HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0. Those results cover six frontier models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4. The release also includes two related cyber evaluations, Cyber CTFs and The Last Ones, using a different, partially overlapping model set. AISI’s accompanying paper studies how benchmark performance depends on inference-time compute and evaluation protocol. In Humanity’s Last Exam, performance changes with both factors: when models received correctness feedback from an oracle after each attempt, they continued solving additional tasks as token use increased. The organizations argue that transcript-level transparency helps researchers diagnose results and understand how setup choices affect reported performance. They invite model developers, evaluation developers, and evaluation, governance, and policy researchers to contribute data, benchmarks, and analyses through the shared infrastructure.