this is truly exceptional work. i had the chance to talk to Steven a bit about how he went about designing and building TB-Science, and i came away really impressed
making benchmarks is easy. making an outstanding one is unbelievably hard. the incentives are stacked against it, from diluting task quality to ship fast, to building hill-climbable problem shapes you can beat by buying envs on the market
see, many folks think of benchmarks as marketing tools, and in a way they are! but the best ones are way more. to me they're basically the AI version of GDP. heavily gamed and debated, yet still indispensable. a good public eval doesn't need to be ungameable. it needs to be hard enough to game that gaming it looks a lot like actually getting better
steven and his team built something special here. measuring the frontier of science and how fast (or how slow!) models are pushing it further is arguably one of the most important things in the world today. doing it with relentless rigor makes it a public good
can't wait to see TBScience v02 come out, and if you're a scientist, you should consider contributing!
We're releasing Terminal-Bench-Science: a benchmark for evaluating AI agents on research workflows across scientific domains.
An ongoing Stanford-led community effort, built by the team behind Terminal-Bench together with scientific domain experts at research institutions worldwide. v0.1 has 70 tasks. Claude Opus 5 solves only ~30%.
1/n 👇