Snorkel AI Commits $30 Million to Expand Open Benchmarks
Summary
Snorkel AI announced that it will expand its Open Benchmarks Grants commitment tenfold, from an initial $3 million to $30 million, to support researchers, domain experts, and open-source teams developing robust, open evaluations for frontier AI. The program will fund new benchmarks and measurement methods for safety and alignment, cybersecurity, physical AI, scientific agents, long-running agent work, open-ended outputs, dynamic settings, and human uplift. Snorkel is also launching an Open Benchmarks Red Team to work with benchmark builders on reward-hacking exploits, data contamination, ambiguous or unsolvable instructions, broken verifiers, stale task sets, and gaps in distributional and task diversity. A new Snorkel Research Fellowship will support independent researchers with a research co-advisor, domain experts, compute, and engineering assistance, while allowing fellows to set their own questions and lead projects through public release. The company says its earlier grants and collaborations supported benchmarks including Terminal-Bench, TB-Science, ARC-AGI-3, OSWorld 2.0, Agents’ Last Exam, Continual Learning Bench, Senior SWE-Bench, and SlopCodeBench. Snorkel argues that benchmark development is lagging behind model development, allowing simple or static tests to become overfit targets, a problem it calls “benchmaxxing.” It proposes broader coverage across environment and input complexity, autonomy horizons, and output complexity. The article also frames open benchmarks as a form of empirical research infrastructure because transparent methods enable inspection, reproduction, adversarial testing, and ongoing maintenance. Snorkel invites benchmark proposals, red-team applications, and fellowship applications.