We are a researcher community developing scientifically grounded research outputs and robust deployment infrastructure for broader impact evaluations.

🚨The @AISecurityInst has teamed up with @evaluatingevals to make AI evaluation results more reproducible! 🚀 Many of UK AISI’s publicly reported evaluation methods and findings will now live on Eval Cards, under the EEE Schema. More details 👇 evalevalai.com/infrastructur…
5
12
54
3,834
✅The release includes verified results, context, and configuration information for five popular evaluation benchmarks including HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro, Terminal-Bench 2.0. You can check out EvalCards at evalcards.evalevalai.com/
7
238
It's time agent evals lived in one place!🏠 We’re launching Every Agent Ever (EAE), building on EveryEvalEver (EEE) to create a shared schema for reporting, storing, and analyzing agent runs. 🚀 Help us shape it! Fill out the form & join the Slack 👇 forms.gle/Ckw3vKSWv94cJvR48
1
6
35
2,861
If you were thinking, oh Every Eval Ever is quiet probably nothing goes out there, well... Data keeps pouring in Oh, and we always look for contributors or people to lead the next projects (more soon) say hi or tag your friend, PI,...
4
9
813
EvalEval Coalition retweeted
Replying to @nardosamengesha
Some of our biggest contributors at @evaluatingevals are undergrads!
2
1
8
974
At EvalEval, we’re working to make AI evaluation: 🔹 More transparent 🔹 More reproducible 🔹 More reusable 🔹 Grounded in community-driven standards With #EveryEvalEver, we’re building a shared infrastructure for AI evaluations. 📖 IBM Research: research.ibm.com/blog/every-… 2/3
3
6
212
It’s a wrap! Thanks to everyone who stopped by our posters for the questions and discussions. And to all the authors: well done! 👏👏
3(!!!) posters down, 1 more main and 1 workshop poster to go :) #ICML2026
19
1,006
EvalEval Coalition retweeted
Attending #ICML2026 @icmlconf in Seoul 🇰🇷 this week? We're presenting two papers tomorrow on AI evaluations as part of the @evaluatingevals coalition. If you're around, come by! 📊 1st paper: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation 📄 arxiv.org/pdf/2602.16763 🕝 2:30 PM KST 📍 Hall A • Poster #4005 TL;DR: Many widely used AI benchmarks are no longer able to reliably distinguish today's frontier models. In this paper, we: • introduce the first uncertainty-aware measure of benchmark saturation, • analyze 60 widely used LLM benchmarks, • find that nearly half already exhibit high saturation, and • identify benchmark design choices associated with longer-lasting discriminative power. 🧵 1/2
2
2
30
1,919
🚨 Are you in San Diego for #ACL2026? So is EvalEval📍 The Second Workshop on Evaluating Evaluations is taking place today, focusing on AI evaluation in practice. ⏰ Time: 2:00 p.m 📍 Location: Harbor A
2
17
927
We will kick it off with a panel on Sociotechnical Perspectives & Implications of Evals featuring @ang3linawang (@Cornell), Wesley Deng (@MSFTResearch), and @sebgehr (Head of Responsible AI at @business), followed by an amazing lineup of spotlight presentations and posters.
1
6
184
We have an amazing program led by the dream team of organizers, @UsmanGohar, Jennifer Mickel, @batznerjan, @LChoshen, Ichhya Pant, Michelle Lin, @akhtarmubashara, @evijit. 🤩 Special thanks to our hosts @huggingface, @EdinburghUni, @AiEleuther.
3
157
EvalEval Coalition retweeted
Spent the morning checking out this platform after watching a great demo from Wm Matthew Kennedy for the @publicainetwork group! It has a lot of potential for making the evals ecosystem easier to navigate and helping technical AI governance rest on a shared common ground.
🚀We launch Evaluation Cards (beta): a centralized public record of AI evaluation results 🚀 Not another leaderboard. Every score comes with who ran it, the settings they used, what the benchmark tests and the other results reported for the same model, side by side. 🧵👇
1
6
312
Interested in the latest research on AI evaluations and heading to #ACL2026 or #ICML2026? ✈️ We went through the accepted papers and curated a selection worth checking out if you’re interested in #AI #evaluations, #benchmarking, evaluation #methodology, and #measurement. 🧵👇
2
9
41
3,079
🚀 Plus our own #EvalEval papers at #ICML2026 @icmlconf: • When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation • Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
1
1
165
Check out the full accepted paper lists: 📚 @aclmeeting 2026: 2026.aclweb.org/program/acce… 📚 @icmlconf 2026: icml.cc/virtual/2026/papers.… Did we miss any interesting papers on AI evaluations? Share your recommendations! ⬇️
1
133