A state AI institute starts publishing its evaluations in a shared schema
The EvalEval Coalition announced on 22 September 2026 that the UK AI Security Institute is using EvalEval's infrastructure to publish its evaluation results openly.
The two have worked together before: research that began at a joint workshop alongside NeurIPS 2025, and feedback from the Institute helped shape the Every Eval Ever (EEE) schema. This phase puts that shared infrastructure into use rather than into a paper.
The problem being addressed is stated plainly in the post. As AI deployment accelerates, evaluations are becoming an important source of evidence about how a model or system performs — yet results are reported across many formats, platforms and outlets, often without enough information to reproduce them, and re-running an evaluation may itself be prohibitively expensive. EvalEval's answer is a shared reporting schema plus an open platform, Evaluation Cards, that puts results and the information needed to interpret them into one structure.

The part worth noticing
The cost sentence is the load-bearing one. In most fields "we could not reproduce it" is a finding. In evaluation work it is frequently not even reachable: the compute to re-run a benchmark can exceed what an independent checker will spend, so the published number goes unchallenged not because it survived scrutiny but because scrutiny was unaffordable. A schema that carries the conditions alongside the score does not make re-running cheap — it makes the claim checkable on its own terms, which is the next best thing.
And the identity of the adopter is the story, more than the schema. A national institute is not a lab publishing a leaderboard; its numbers are the sort that get quoted into policy. When that body commits to a structured, open format, the format stops being a community convention and starts being something other publishers get asked why they are not using.
What this does not establish is that the results themselves become comparable. A shared schema standardises how an evaluation is described, not what it measures — two cards can be complete, well-formed and still be measuring different things. That is a real limit, and the post does not claim otherwise.