🚨The @AISecurityInst has teamed up with @evaluatingevals to make AI evaluation results more reproducible! 🚀
Many of UK AISI’s publicly reported evaluation methods and findings will now live on Eval Cards, under the EEE Schema. More details 👇
evalevalai.com/infrastructur…
✅The release includes verified results, context, and configuration information for five popular evaluation benchmarks including HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro, Terminal-Bench 2.0.
You can check out EvalCards at evalcards.evalevalai.com/
It's time agent evals lived in one place!🏠
We’re launching Every Agent Ever (EAE), building on EveryEvalEver (EEE) to create a shared schema for reporting, storing, and analyzing agent runs. 🚀
Help us shape it! Fill out the form & join the Slack 👇
forms.gle/Ckw3vKSWv94cJvR48
If you were thinking, oh Every Eval Ever is quiet probably nothing goes out there, well... Data keeps pouring in
Oh, and we always look for contributors or people to lead the next projects (more soon)
say hi or tag your friend, PI,...
🚀 Exciting to see Evaluating Evaluations in the media!
Over the past weeks, our initiative has been featured by both @IBMResearch and @mozilla, highlighting the importance of evaluations for measuring AI progress.
📖 Mozilla: blog.mozilla.ai/from-evaluat…
1/3
At EvalEval, we’re working to make AI evaluation:
🔹 More transparent
🔹 More reproducible
🔹 More reusable
🔹 Grounded in community-driven standards
With #EveryEvalEver, we’re building a shared infrastructure for AI evaluations.
📖 IBM Research: research.ibm.com/blog/every-…
2/3
A huge thank you to everyone contributing to this growing ecosystem! 🙌
If you’d like to learn more about the EveryEvalEver schema:
📄 Paper: arxiv.org/abs/2606.14516
3/3
Attending #ICML2026@icmlconf in Seoul 🇰🇷 this week?
We're presenting two papers tomorrow on AI evaluations as part of the @evaluatingevals coalition. If you're around, come by!
📊 1st paper: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
📄 arxiv.org/pdf/2602.16763
🕝 2:30 PM KST
📍 Hall A • Poster #4005
TL;DR: Many widely used AI benchmarks are no longer able to reliably distinguish today's frontier models.
In this paper, we:
• introduce the first uncertainty-aware measure of benchmark saturation,
• analyze 60 widely used LLM benchmarks,
• find that nearly half already exhibit high saturation, and
• identify benchmark design choices associated with longer-lasting discriminative power.
🧵 1/2
🚨 Are you in San Diego for #ACL2026? So is EvalEval📍
The Second Workshop on Evaluating Evaluations is taking place today, focusing on AI evaluation in practice.
⏰ Time: 2:00 p.m
📍 Location: Harbor A
We will kick it off with a panel on Sociotechnical Perspectives & Implications of Evals featuring @ang3linawang (@Cornell), Wesley Deng (@MSFTResearch), and @sebgehr (Head of Responsible AI at @business), followed by an amazing lineup of spotlight presentations and posters.
Spent the morning checking out this platform after watching a great demo from Wm Matthew Kennedy for the @publicainetwork group! It has a lot of potential for making the evals ecosystem easier to navigate and helping technical AI governance rest on a shared common ground.
🚀We launch Evaluation Cards (beta): a centralized public record of AI evaluation results 🚀
Not another leaderboard. Every score comes with who ran it, the settings they used, what the benchmark tests and the other results reported for the same model, side by side. 🧵👇
🚀 Plus our own #EvalEval papers at #ICML2026@icmlconf:
• When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
• Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting