terminalbench retweeted
Before their first release, the @terminalbench team called off a task-writing meetup, figuring there was no way to get more than ten people in a room to talk about the project. Tonight we squeezed ~150 into Laude Lab! The team covered the state of the bench and @harborframework, the process behind Terminal-Bench-Science, and how to keep pace as models get better, faster: continuous benchmarks, real-world evals, and long-horizon multi-agent challenges. @alexgshaw @ryan_marten @StevenDillmann @Mike_A_Merrill @andykonwinski
5
10
118
12,003
terminalbench retweeted
Terminal-Bench meetup tonight *with* swag. Register if you haven’t already! Benchmarks, RL environments, new features, and state of the union luma.com/tbench
17
10
115
11,213
terminalbench retweeted
fast track to get model labs to care about the capabilities you care about: contribute a task to Terminal-Bench if you have built a benchmark around a specific use case, DM me and we can collaborate on a TB task for the next release
if you are building a product using AI, you should be spending >25% of your time making benchmarks and trying to get the model labs to care about said benchmarks easiest path to accelerate your progress as a company
7
4
62
6,528
We're hosting a meetup! Come meet the community behind the benchmark, celebrate recent releases, and join our discussions about the future of agent evaluation. luma.com/tbench
3
5
54
20,582
terminalbench retweeted
(another) new SOTA on @terminalbench!
1
6
48
3,120
terminalbench retweeted
Terminal-Bench-Science 0.1 is the #1 featured benchmark on @AnthropicAI’s new Claude Fable release 🚀 terminal-bench-science.ai
Replying to @claudeai
Across our benchmarks, the model sets a new standard. It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5. On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5.
13
20
136
8,965
terminalbench retweeted
Across our benchmarks, the model sets a new standard. It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5. On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5.
128
449
5,503
1,183,744
terminalbench retweeted
New SOTA on @terminalbench!
We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work.
3
6
67
3,619
Terminal-Bench 4.0 out now!
We've pushed a version update to the Terminal-Bench dataset and leaderboard. Terminal-Bench 4.0 calibrates task resources (time, cpu, memory), implements task fixes, and removes saturated tasks.
3
1
17
1,806
Announcing Terminal-Bench-Science!
We're releasing Terminal-Bench-Science: a benchmark for evaluating AI agents on research workflows across scientific domains. An ongoing Stanford-led community effort, built by the team behind Terminal-Bench together with scientific domain experts at research institutions worldwide. v0.1 has 70 tasks. Claude Opus 5 solves only ~30%. 1/n 👇
2
24
1,545
terminalbench retweeted
Congrats to Z.ai for the strong performance on Terminal-Bench 3.0! One of the biggest pieces of feedback we have gotten for TB3 is to increase the timeouts. We calibrated timeouts against frontier models during development, but inference speed can still be a confounder on some of the tasks. Terminal-Bench numbers on the GLM-5.3 model card are reported with increased timeouts (likely for this reason). Look out for Terminal-Bench 4.0 releasing soon with increased timeouts, other task improvements, and a handful of new tasks.
Introducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model - A major leap in cybersecurity, setting a new standard among open models Tech Blog: z.ai/blog/glm-5.3
9
3
65
5,283
terminalbench retweeted
Grok 4.6 now on the Terminal-Bench 3.0 Leaderboard!
6
9
183
10,957
terminalbench retweeted
Grok 4.5 is SOTA on TB2.1... at reward hacking In all seriousness, even after zeroing out reward hacks, it is #4 on the TB2.1 leaderboard and lands on the Pareto for both cost and speed. (charts and reward hacking links in 🧵)
4
5
83
6,810
terminalbench retweeted
GPT‑5.6 Sol sets a new state of the art on Terminal‑Bench 2.1, which tests complex command-line workflows requiring planning, iteration, and tool coordination.
108
233
3,940
1,926,937
terminalbench retweeted
Can agents build complete projects that deliver real value? We’re launching Terminal Bench Challenges: 3 unsolved tasks which could make a real impact on the open source community if solved. These tasks provide a testing ground for optimizations both on the model and harness level on our continuous leaderboard for each task.
5
11
39
5,006
Introducing Terminal-Bench Challenges! A new capability has emerged at the frontier: agents completing large-scale projects autonomously. To test this capability, we felt another flavor of benchmark was needed. Terminal-Bench Challenges are long-horizon, token-intensive, single-task benchmarks. Today we are releasing our first 3 challenges.
3
13
50
6,809
Terminal-Bench Challenges is inspired by previous projects exploring long-running agents including Carlini's C compiler and Cursor's browser. Join the effort! If you have ideas for further challenges, come hang out in the tb-challenges discord channel discord.com/invite/2Pe5uWGcV…
3
285