A team of researchers has proposed Benchy, a formal language and execution engine meant to make AI benchmarks precise enough for machines to check, not just humans to eyeball.
Benchy defines a benchmark as three fixed pieces: a program, a scoring function, and a dataset, kept separate from whichever AI system takes the test. Benchmarks are written in a canonical YAML format, tagged against a shared task/domain/language ontology, then compiled into a JSON intermediate representation that the engine actually runs. That compilation step only changes format, the paper stresses, not meaning: it won't quietly patch a broken benchmark definition or fill in defaults you didn't ask for. Every AI system plugs into the engine through one fixed contract, a named-field input in, a named-field output out, so the messy work of hooking up a model never leaks into how the benchmark is scored.
The pitch here is less about any single benchmark and more about the format wars underneath all of them. Right now, benchmark suites are typically bespoke code, which makes results hard to audit and easy to game with hidden preprocessing. A shared, machine-checkable spec would let anyone inspect exactly what a benchmark measures and how a score gets computed, rather than trusting a leaderboard's word for it.
Whether Benchy becomes that shared spec or just another one competing for attention is the open question. The paper itself only commits to the engineering contract for a first engine implementation, which is a modest way of saying: this is a proposal, not yet a standard.