Over a year ago, we released SWE-Bench Pro.
We refreshed the benchmark today to improve its quality based on community feedback. Our update is meant to ensure the benchmark is clean and accurately measures model capabilities on SWE-Bench Pro tasks.
In our update, we observe a ~20% performance drop on our held out private example set of 272 tasks compared to the public leaderboard. We further identify a HARD subset of 51 tasks that is discriminative of the frontier model performances from the rest of the pack.
Refreshing benchmarks, and sharing what we learn along the way, does more to advance model evaluation than deprecating them outright.
That's why we're sharing this updated version: it addresses issues we identified over time, and we want the broader community to benefit from those findings.
SWE-Bench Pro V2 is live.
What’s new: 🧵