Benchy Proposes a Universal Language for Task-Oriented AI Benchmarks
Summary
Benchy is a semantic language and execution engine for benchmarking AI programs. It defines a benchmark as a program, scoring function, and dataset, while a run binds that benchmark to an AI system. Benchmarks are authored in canonical YAML, where each semantic concept has one valid syntax, and are classified through a shared task, domain, and language ontology. The YAML is deterministically compiled into canonical JSON, with compilation changing representation but neither repairing invalid definitions nor adding hidden defaults. Programs use fixed schemas of named input and output fields, and leaf output fields define the scoring dimensions. Benchy exposes a universal runtime contract consisting of a named-field input object and a named-field output object, allowing external AI systems to adapt at the boundary without changing benchmark semantics. The paper describes the object model, validation rule, scoring and failure semantics, compilation and execution architecture, and the scope of the current language. An appendix specifies the normative engineering contract for the first engine implementation.