How UK AISI and EvalEval Are Making Benchmark Results Reproducible
By Jakub Antkiewicz
•2026-09-23T13:22:57Z
A New Standard for AI Benchmarking
The UK AI Safety Institute (AISI) has partnered with the EvalEval Coalition to openly publish its AI model evaluation results, a significant step toward addressing the industry's pervasive issue with benchmark reproducibility. This collaboration leverages EvalEval's infrastructure to create a common, verifiable format for reporting model performance. The initiative aims to solve a critical problem where results are often fragmented across different platforms and lack the necessary details for independent verification or direct comparison.
The Data and Infrastructure
The initial release from AISI accompanies its research paper, "How Inference Compute Shapes Frontier LLM Evaluation," and makes key data available through EvalEval's "Evaluation Cards" platform. This system uses the shared "Every Eval Ever" (EEE) schema, which has been refined with feedback from the Institute. The release provides transcript-level transparency and detailed configuration information, allowing others to see precisely how evaluation protocols and compute allocation can influence outcomes.
- Benchmarks Released: HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0.
- Models Covered: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4.
- Additional Evaluations: The release also includes results from two cyber-focused evaluations, Cyber CTFs and The Last Ones.
Impact on the AI Ecosystem
By publishing verified reference points, AISI and EvalEval are establishing a new baseline for transparency in the AI ecosystem. This enables researchers and developers to conduct more reliable meta-research and better understand how specific setup choices affect performance. The adoption of the EEE schema by a major government-backed institute encourages other model developers, evaluators, and governance bodies to move away from opaque leaderboards towards a shared, scientifically rigorous standard for reporting AI capabilities.
Strategic Takeaway: The formal collaboration between a state-backed entity like the UK's AI Safety Institute and a research community like EvalEval institutionalizes the push for evaluation transparency. This shifts the industry's focus from leaderboard scores alone to the underlying scientific reproducibility, potentially making detailed, verifiable evaluation reporting a new competitive standard.