Eval scores are hard to interpret and compare when the setup that produced them isn't reported. We've been working with @evaluatingevals on standardising eval reporting so results are more reproducible and transparent. As a first step, we've contributed verified results to their platform. 🧵
🚨The @AISecurityInst has teamed up with @evaluatingevals to make AI evaluation results more reproducible! 🚀 Many of UK AISI’s publicly reported evaluation methods and findings will now live on Eval Cards, under the EEE Schema. More details 👇 evalevalai.com/infrastructur…

Sep 22, 2026 · 5:41 PM UTC

2
3
32
2,681
Even when eval results are public, they’re often scattered across formats and platforms, with too little detail to reproduce them. EveryEvalEver provides a common reporting schema; Evaluation Cards brings results together in one place for easier comparison.
1
3
121
Our first contribution covers verified results, contexts, and setup config across several public benchmarks (HealthBench, FrontierMath, HLE, SWE-Bench Pro, Terminal-Bench 2.0) and six frontier models, drawn from our recent work on test-time compute scaling. aisi.gov.uk/blog/more-comput…
2
4
108
Openly reported setup lets others assess how configuration choices influence measured performance. If you develop models or evals, we'd encourage you to report your results in EEE and dig into the open comparisons on Evaluation Cards! evalevalai.com/infrastructur…
2
44
Sort replies: Relevant Recent Liked