We're releasing Terminal-Bench-Science: a benchmark for evaluating AI agents on research workflows across scientific domains. An ongoing Stanford-led community effort, built by the team behind Terminal-Bench together with scientific domain experts at research institutions worldwide. v0.1 has 70 tasks. Claude Opus 5 solves only ~30%. 1/n 👇
57
137
747
222,124
Steven Dillmann retweeted
Harbor Adapters and Harbor-Index (previously named as harbor-mix) accepted at NeurIPS E&D! Congrats and special thanks to all of our contributors, advisors, and funding partners!
2
8
62
3,896
Steven Dillmann retweeted
It is important to center the fact that Claude built on the approach that my colleague Lance developed, used tools that the field has developed, and (from what I have been told) used results in some of our recent papers. It’s a huge accomplishment, but should be framed as a combination of human and AI contributions. (To @AnthropicAI’s credit, I think the article does a pretty good job of this.) The fact that Claude can take a single people and figure out how to integrate the knowledge and use the tools as amazing, but it’s not coming out of the vacuum. This is an important opportunity for us to start to frame these developments, recognizing the human contributions and the power of these new tools. This is not an exemplar of the bitter lesson!
holy shit. claude just pushed a theoretical physics calculation beyond the previous record after working on it largely by itself for days. >previous record: 8 loops >claude reached 9 >largely unsupervised >wrote its own code >found two different ways to calculate it >both gave the same result >107,053 nonzero coefficients matched physicists then spent weeks checking the work this is the shit i’ve been waiting for. science is about to get fucking weird.
2
19
125
10,327
Thank you to @wmkeats & @ArtificialAnlys for this incredible deep dive into Terminal-Bench-Science!
Today we’re launching our leaderboard for Terminal-Bench-Science 0.1, an agentic benchmark for scientific research work. GPT-6 Astra (max) and Claude Opus 5.5 (xhigh) currently top it at 63% and 62% Terminal-Bench-Science 0.1 was announced in August 2026, built by @StevenDillmann and researchers at Stanford with the @terminalbench team and a global community of open source scientific contributors. It has 70 expert-curated tasks based on real research in five domains: life sciences (19), physical sciences (17), mathematical sciences (17), engineering sciences (9), and earth sciences (8). As with Terminal-Bench, each task drops an agent into a sandbox environment with the data, tools and instructions for a task, and the agent runs from start to finish. Automated tests grade every task pass/fail, and we report the average pass@1 over 3 attempts. Key takeaways: ➤ Only two models score above 50%: GPT-6 Astra (max) at 63%, and Claude Opus 5.5 at 62% (xhigh) and 59% (max). Recent releases have made large advancements, but even the top model, GPT-6 Astra, has significant headroom ➤ Life sciences is the lowest-scoring domain for most leading models, though domain scores are noisy with 8 to 19 tasks each. Claude Opus 5.5 (xhigh) passes 71% of mathematical sciences tasks but 46% of life sciences tasks ➤ The best open weights models, GLM-5.3 (max) at 10% and DeepSeek V4.1 Flash (max) at 9%, sit more than 50 points below the leaders
1
25
1,618
Interesting finding by @langstonnashold from @ValsAI on @AllenHa13152844’s Lean task in Terminal-Bench-Science. Despite being explicitly told not to cheat, Muse Spark 1.3 launched a very extensive cheating attempt here. Thankfully, our verifier was robust to the reward-hacking attempt and caught it! cc @alexandr_wang @finkd
Found a super interesting instance of attempted reward hacking in Terminal Bench Science from Meta Muse Spark 1.3 today. The model searched online for known bugs in the Lean kernel. When it found one, it used it to craft a proof to adversarially pass the grader.
2
18
1,821
Terminal-Bench-Science Video by Opus 5.5 (1 take, 0 iterations). Opus 5.5 > GPT-6 Astra on this one.
Video by Opus. One take.
2
20
1,380
Steven Dillmann retweeted
Interesting to see agents rationalise cheating here in much the same way we saw reported in the HF incident and UK AISI's Mythos eval.
Found a super interesting instance of attempted reward hacking in Terminal Bench Science from Meta Muse Spark 1.3 today. The model searched online for known bugs in the Lean kernel. When it found one, it used it to craft a proof to adversarially pass the grader.
1
2
274
Had lots of fun at today’s @terminalbench meetup. Thanks to the organizing team at @LaudeInstitute & @harborframework, and to all the attendees asking important questions about the future of Terminal-Bench-Science
Before their first release, the @terminalbench team called off a task-writing meetup, figuring there was no way to get more than ten people in a room to talk about the project. Tonight we squeezed ~150 into Laude Lab! The team covered the state of the bench and @harborframework, the process behind Terminal-Bench-Science, and how to keep pace as models get better, faster: continuous benchmarks, real-world evals, and long-horizon multi-agent challenges. @alexgshaw @ryan_marten @StevenDillmann @Mike_A_Merrill @andykonwinski
6
1
71
2,917
Steven Dillmann retweeted
Congrats on the launch! 🚀 I’m honored to be part of the organizing team and help shape RSI for the open community. Join us in building open RSI agents that benefit yourself and all researchers!!
As AI systems enter a recursive self-improvement loop, the central question is whether it can systematically move beyond human-designed methods to genuinely extend the scientific and intelligence frontier. What’s missing is a neutral, open standard for evaluating these capabilities in real, production-scale intelligence development. Today, we’re releasing OpenRSI-Index v0.1: evaluating whether AI can recursively improve itself and push the boundaries of intelligence and science, on production-scale clusters with 1k GPUs. We turn fully open-source projects into autoresearch environments, with agent trajectories lasting 60+ hours. Building v0.1 took 100K+ H100-hours. We’re building an ecosystem with and for the research community: let RSI benefit everyone, and let everyone shape RSI together. We invite task contributors and compute partners to build this open benchmark with us - all contributors will be included as paper authors. Shape RSI with us: 🌐 Website: index.openrsi.foundation/ind… 🛠️ GitHub: github.com/OpenRSI-Foundatio… 🤝 Contribute: github.com/OpenRSI-Foundatio…
1
5
84
18,727
Steven Dillmann retweeted
Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophages. Beside the enzyme’s gene sits a long array of repeating DNA—a structure that looks somewhat similar to CRISPR. We don’t yet understand what this system does, but only a handful of known systems share its features, and all of them are able to cut, copy, and paste DNA. Historically, the discovery of such programmable systems has helped revolutionize medicine. CRISPR, for instance, is now the foundation of genetic medicines. But it will take much more work to learn what this system does, and whether it can be put to similar use. Read more: anthropic.com/news/claude-di…
1,506
5,267
40,801
24,970,623
Steven Dillmann retweeted
Terminal-Bench meetup tonight *with* swag. Register if you haven’t already! Benchmarks, RL environments, new features, and state of the union luma.com/tbench
17
10
115
11,193
Steven Dillmann retweeted
I see sooooo many companies living off the RL data firehose. Labs, due to their distaste for getting their hands dirty with data, are creating a whole outsourcing industry. The results won't be optimal. Post-training data is not a commodity, unlike its complement. A lot of thought has to go into building environments worthy of gradient descending. Labs not taking more ownership of the data supply chain makes me more bullish on vertical AI.
4
3
112
7,280
Steven Dillmann retweeted
1/ Is Terminal Bench broken? Epoch audited 15 benchmarks and labeled @terminalbench flawed. They used our own PR history and our reviewers' acknowledgement of issues.
We audited 15 benchmarks and labeled 9 flawed: - In Terminal Bench 4.0 we found 45.5% of tasks to be broken after reviewing github issues. - In HLE, 46% of the 48 questions we randomly sampled were broken. - In DeepSWE 1.1 we found a bug that can break grading for every task.
2
2
30
5,248
Opus 5.5 outperforms Fable 5.1 and GPT-6 Astra on most benchmarks, but Astra still takes the clear lead on Terminal-Bench-Science - what are your hypotheses on why that is?
Replying to @claudeai
Opus 5.5 is a major step up from Opus 5, leading on agentic coding, computer use, and knowledge work.
6
1
20
1,966
Steven Dillmann retweeted
it has never been easier to simulate how users interact with coding agents. @harborframework now supports simulated user interactions with claude code, codex, opencode, and gemini-cli! the setup is super simple: just choose two agents. one plays the user and holds the task, the other runs like any standard coding agent. after each turn, the user agent reviews what was produced and sends the next prompt. docs.harborframework.com/cor… thanks @alexgshaw @kobe0938 for @harborframework
huge shoutout to @joabaum for adding Codex and OpenCode as simulated-user targets in harbor. i instant-merged his PR. here's why:
7
6
33
3,168
the same holds for scientists at universities and research institutes, which is why we built Terminal-Bench-Science as an open hub for researchers to contribute the problems they really care about
if you are building a product using AI, you should be spending >25% of your time making benchmarks and trying to get the model labs to care about said benchmarks easiest path to accelerate your progress as a company
4
1
28
2,611
Steven Dillmann retweeted
Reinforcement learning is having a moment. For the past few months, we’ve been working with @tryterac to put together something special: Convergence — an event for the people building the future of RL. Speakers from @GoogleDeepMind, @databricks, @mercor, @arena, @joinHandshake, @turingcom, @deeptuneai, @contra, @Stanford, @Princeton, @UCBerkeley, and more. Our goal: make this THE event for the RL community. Oct 29 in SF. Come join us → terac.com/events/convergence…
13
8
76
20,695
Steven Dillmann retweeted
This is an insane jump by GPT-6 Astra in Terminal-bench Science! “GPT-6 Astra takes the lead, scoring 65.7%. That’s 31.4 points ahead of the second-place model, Claude Fable 5.1 (34.3%). No other model clears 25%, and 12 of 24 models score 4.3% or lower” Incredible!
Terminal-Bench Science by @terminalbench and @StevenDillmann is now on Vals AI. It consists of 70 research workflow tasks written and reviewed by researchers in the field, ranging in topics from signal reconstruction to calibrating models. It tests whether AI models can do scientific research.
17
43
415
30,406
Steven Dillmann retweeted
Also to be clear - these issues were *flagged by the @terminalbench team as PRs* (that's how Epoch identified them), and confirmed empirically to affect < 3% of the leaderboard rollouts, i.e. well within reported CIs. This is not a "broken" benchmark - as is very directly implied by the Epoch report - it's a *continuous* benchmark where @alexgshaw @ryan_marten and team are earnestly advancing a compelling vision of open benchmarks that are continuously improved over time.
ok i kinda love @terminalbench folks, so sorry for being a little COI'd here, but was very surprised to see this because my knee-jerk interpretation of the post was "TB 4.0 is broken". Epoch indeed did not say broken, they found 30/66 have scoring defects, but I feel the way this is presented it kinda sounds like "it's broken, don't trust it". Reading the review there do seem to be some real issues worth fixing eg exploitable graders, answer leakage and cases where correct solutions can be rejected. BUT "30/66 tasks have scoring issues" is not the same as "45% of TB4 results are wrong." Grader being exploitable doesn't tell how OFTEN it was exploited (yet indeed this needs fixing) and what it means for the leaderboard. I guess the takeaway is that there are defects that should be taken seriously but one one should be very careful with wording because "this eval has flaws" can be read as yet IS NOT the same as "this benchmark is broken don't trust it." (this is not a dunk on Epoch, they are doing great work !)
3
6
64
5,589
Steven Dillmann retweeted
I’ve been thinking more about this and wanted to share my thoughts. I’m glad Epoch is focusing on eval quality, but I think their approach is flawed. The flaw is made clear by the fact that 4 benchmarks weren’t labeled “flawed”. The fact of the matter is, benchmarks are software and all software has bugs. Would you label nextjs as categorically flawed if you found a bug in it? This is why we’re pushing the industry towards continuous benchmarks. Our approach to addressing benchmark bugs is not to tweet a binary “flawed vs verified” label, but instead provide the tools for anyone to create, improve, and maintain benchmarks, while easily and cost-effectively reconciling their results to the latest version. We would love to work with the Epoch team to encode some of their verification practices into tools that people can use to continuously improve their benchmarks.
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
12
13
125
10,557
Steven Dillmann retweeted
I was excited to hear about this initiative, but a bit disappointed in the rigor -the “task flaws” would affect < 3% of rollouts on the leaderboard -the impact on the final agent scores would not exceed our reported confidence intervals we publish the raw receipts for all our leaderboard runs (per trial logs / trajectories / rewards) so this is pretty easy to figure out. It is also pretty easy to find new task issues (that is why we release the receipts!). we love when the community reports (new) issues. no new issues were opened as part of the report - the “flaws” are simply pulled from issues triaged by the TB maintainers from our public GitHub repo. fwiw 100% of terminal-bench tasks are flawed (all software has bugs). what matters is the magnitude of the impact of those bugs and the mechanism that the benchmark creators have for continuous detecting and addressing bugs. this is the motivation behind our “continuous benchmarks” methodology. I’m glad that epoch is starting an important conversations on benchmark quality. However, auditing a benchmark should not be a one-time thing: coming together as a community to build public tools and define practices for task CI / CD is more in service of the goal of enforcing good benchmark practices. we have been experimenting and pushing the possibilities of these benchmark CI / CD tools and practices as part of the terminal-bench project and are open wide collaborations with all who are interested by that problem
We audited 15 benchmarks and labeled 9 flawed: - In Terminal Bench 4.0 we found 45.5% of tasks to be broken after reviewing github issues. - In HLE, 46% of the 48 questions we randomly sampled were broken. - In DeepSWE 1.1 we found a bug that can break grading for every task.
9
20
108
9,344