How UK AISI and EvalEval Are Making Benchmark Results Reproducible
Blog post from Hugging Face
The EvalEval Coalition and the UK AI Security Institute (AISI) are expanding their collaboration to make AI benchmark reporting more reproducible, transparent, and easier to verify through EvalEval’s Every Eval Ever schema and Evaluation Cards platform. AISI has released verified evaluation methods, configurations, contextual information, and results from research on how inference-time compute and evaluation protocols affect frontier model performance, covering benchmarks including HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0, as well as related cyber evaluations. The results span several frontier models, including Claude Opus variants and GPT-5 variants, and illustrate how factors such as token budgets and correctness feedback can substantially affect reported scores. The initiative aims to provide reference points for comparing studies conducted under different conditions, support meta-research into evaluation practices, and encourage model developers, benchmark creators, and policy researchers to use common reporting standards.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Guardrails | 4 | 35 | 22 | 12 | -94% |
| LLM | 3 | 747 | 162 | 79 | -85% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.