NeurIPS 2025 Workshop. Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling

San Diego
We are happy to announce our @NeurIPSConf workshop on LLM evaluations! Mastering LLM evaluation is no longer optional -- it's fundamental to building reliable models. We'll tackle the field's most pressing evaluation challenges. For details: sites.google.com/corp/view/l…. 1/3
3
3
36
29,262
LLM Evals Workshop @NeurIPS retweeted
It has been a super fun day @LLM_eval workshop @NeurIPSConf with amazing talks, posters, and an engaging panel discussion! @dawnsongtweets @natolambert @orf_bnw @sanmikoyejo @abeirami @hamishivi @MariusHobbhahn @beyzaermis @Diyi_Yang @attaluri_nithya @RishiBommasani @YangjunR
6
10
137
17,494
LLM Evals Workshop @NeurIPS retweeted
Our next talk @LLM_eval workshop is by @sanmikoyejo! Upper Level Room 2 @NeurIPSConf
5
24
3,181
LLM Evals Workshop @NeurIPS retweeted
“Good researchers obsess over evals” by @natolambert @LLM_eval workshop!
6
59
4,888
LLM Evals Workshop @NeurIPS retweeted
Bringing the hot take culture to NeurIPS - great talk @orf_bnw!!
Replying to @LLM_eval
@LLM_eval workshop has started with Orhan Firat’s talk at Upper Level Room 2. @NeurIPSConf
1
16
1,928
LLM Evals Workshop @NeurIPS retweeted
Replying to @dawnsongtweets
@dawnsongtweets is giving a talk on agentic evals @LLM_eval workshop!
1
25
1,531
LLM Evals Workshop @NeurIPS retweeted
Replying to @LLM_eval
@LLM_eval workshop has started with Orhan Firat’s talk at Upper Level Room 2. @NeurIPSConf
2
3
33
4,644
LLM Evals Workshop @NeurIPS retweeted
Good researchers obsess over evals The story of Olmo 3 (post-training), told through evals NeurIPS Talk tomorrow. Upper Level Room 2, 10:35AM.
11
46
596
57,469
LLM Evals Workshop @NeurIPS retweeted
I’ll be @NeurIPSConf all week and would love to connect on LLM data, evaluation, benchmarking, and scaling laws. If you’re working on related problems, feel free to reach out. PS: Don’t miss our one-of-a-kind workshop on LLM evaluation: sites.google.com/view/llm-ev…
6
5
97
9,405
🚀 We are thrilled to announce that the LLM Eval Workshop @NeurIPSConf received 244 excellent submissions! 188 papers will be presented in poster sessions, and 5 exceptional works have been selected for oral talks. Check out the accepted papers: sites.google.com/view/llm-ev… 🧵👇
1
5
642
- "The Measure of All Measures: Quantifying LLM Benchmark Quality" -- Jihan Yao, Peter Jin, Ke Bao, Qiaolin Yu et al. openreview.net/forum?id=HpnG…
1
4
759
See you in San Diego on December 7th!
2
150
LLM Evals Workshop @NeurIPS retweeted
I will present my #EMNLP2025 paper at the #NeurIPS2025 LLM Eval Workshop @LLM_eval (Dec. 7th 11:15 - 12:15Poster Session 2). If you are interested in reliable LLM-as-a-judge, please come say hi! ☕️ #AI #LLM #LLMJudge #LLMEvaluation #ConformalPrediction
🤩My FIRST paper received #EMNLP2025 SAC Highlights: "Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal Prediction" Huge thanks to my advisor @jiank_uiuc and collaborators Xinyi Liu, @hangfeng_he , & @jieyuzhao11 ! #AI #NLP #LLM #ConformalPrediction
2
13
1,044
LLM Evals Workshop @NeurIPS retweeted
(1/5) My work, “LLMs Show Surface-Form Brittleness Under Paraphrase Stress Tests”, has been accepted for a contributed talk at @NeurIPSConf 2025 Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling workshop @LLM_eval #NeurIPS #LLM #Evaluation #Robustness #AI #ML
5
6
23
8,548
LLM Evals Workshop @NeurIPS retweeted
Sketched on a few Parisian summer nights with a friend, @ChrisInterno . If you care about (causal) identification in a semi-synthetic future, we’d value your read and critique. Preprint: arxiv.org/pdf/2509.17999 Accepted at @LLM_eval workshop @NeurIPSConf
1
4
261
LLM Evals Workshop @NeurIPS retweeted
The Narcissus Hypothesis: --Recursive training on semi-synthetic corpora enforcing human alignment induces a Social Desirability Bias: world-models (Narcissus) aim to please rather than represent, polluting data lakes and charming us (Echo) into hanging on their every word.
1
4
7
1,124