Epoch AI Lists 85 Benchmarks for Evaluating AI Capabilities
Summary
Epoch AI’s benchmark registry contains 85 evaluations covering mathematics, software engineering, agentic workflows, games, long-context tasks, multimodal systems, science, language understanding, and other areas. The directory can be filtered by creator, review verdict, score range, number of evaluations, and domain; its review categories include verified, flawed, and not enough information. Epoch administers 14 of the listed benchmarks, while 52 are administered by benchmark creators and 18 by model developers. The registry includes the Epoch Capabilities Index, which aggregates multiple benchmarks into a general capability scale, and FrontierMath, a collection of difficult or unsolved mathematics problems. Other entries include MirrorCode, which tests whether models can reimplement complete programs without seeing the original source, and EBR-bench, which measures score improvement across repeated playthroughs of Earthborne Rangers. The page also lists games and knowledge evaluations such as Mystery Game Puzzles, Chess Puzzles, and the verified 1,000-question SimpleQA benchmark. Listed highest scores vary substantially by benchmark, including 3% for FrontierMath Erdős, 73% for MirrorCode, 76% for EBR-bench and SimpleQA, and 98% for FrontierMath Tier 4 v2.