ICML 2026 is almost here, and we're headed to Seoul with 7 accepted papers. 🇰🇷
Our research spans agent training, scientific reasoning, and evaluations advancing the next generation of AI systems.
Here's an early look at the work we'll be presenting 🧵
1
4
32
2,652
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
arxiv.org/abs/2509.16941
1
1
6
398
Online Rubrics Elicitation from Pairwise Comparisons
arxiv.org/abs/2510.07284
1
3
233
Coverage, Not Averages: Semantic Stratification for Trustworthy Retrieval Evaluation
arxiv.org/abs/2604.20763
1
1
158
VeRO: An Evaluation Harness for Agents to Optimize Agents
arxiv.org/abs/2602.22480
1
2
204
SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?
arxiv.org/abs/2604.10718
1
1
147
Imitation Learning for Multi-turn LM Agents via On-policy Expert Corrections
arxiv.org/abs/2512.14895
Jul 2, 2026 · 7:18 PM UTC
1
1
3
3,175
RubricRobustness: Evaluating the Sensitivity of Rubrics-Based Benchmarks to Simple Perturbations
openreview.net/forum?id=2Y3I…
1
1
414







