ICML 2026 is almost here, and we're headed to Seoul with 7 accepted papers. 🇰🇷 Our research spans agent training, scientific reasoning, and evaluations advancing the next generation of AI systems. Here's an early look at the work we'll be presenting 🧵
1
4
32
2,652
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? arxiv.org/abs/2509.16941
1
1
6
398
Online Rubrics Elicitation from Pairwise Comparisons arxiv.org/abs/2510.07284
1
3
233
Coverage, Not Averages: Semantic Stratification for Trustworthy Retrieval Evaluation arxiv.org/abs/2604.20763
1
1
158
VeRO: An Evaluation Harness for Agents to Optimize Agents arxiv.org/abs/2602.22480
1
2
204
SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? arxiv.org/abs/2604.10718
1
1
147
Imitation Learning for Multi-turn LM Agents via On-policy Expert Corrections arxiv.org/abs/2512.14895

Jul 2, 2026 · 7:18 PM UTC

1
1
3
3,175
RubricRobustness: Evaluating the Sensitivity of Rubrics-Based Benchmarks to Simple Perturbations openreview.net/forum?id=2Y3I…
1
1
414
Sort replies: Relevant Recent Liked