Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.

Lake Tahoe
With the new frontier continuing to move quickly, and results from GPT-6 Astra showing clear saturation across a number of benchmarks, we are making some pretty exciting moves at VulcanBench. First, our new eval suite, that we just released will be called VulcanBench SWE v4. I am already seeing it give sub 90% scores for Fable 5.1, and are sure hoping it can challenge Astra too. As usual, all real coding tasks, from real repos. Updates to the benchmarking process also that make it very hard for the model to cheat, gives more opportunities for partial credit, and increases the timeout per task to a flat 10 hours to make sure that if something fails, it鈥檚 not because it ran out of time. Second, with models like Astra and Fable 5.1 out, I think its time to launch a multi-benchmark index, focused 100% on coding, and sourcing some of the coding benchmarks that still stand up to the current frontier. I鈥檒l be calling this the Agentic Coding Capability Index, and it will consist of five different benchmarks including the brand new Terminal Bench 4.0. While VulcanBench will continue dedicated to our mission of building eval suites and benchmarking across effort levels, I am excited to offer more data that can help engineers and engineering leaders make decisions. Live long and benchmark 馃枛 Below is a table showing the benchmarks and weights for the Agentic Coding Capability Index v1:
1
19
5,709
We have decided to stop work on our eval suite for System One Models. Have a ton of respect for Diogo, and we're so busy benchmarking LLMs, if he doesn't want independent benchmarks to test models like Jev, we won't. That being said, we love Jev, it's a great model, Diogo is a freakin genius and built a great team, highly recommend people try it for themselves and see how it might make its way into their workflow. Onward, live long and benchmark 馃枛
to re-iterate, I'm extremely anti-benchmarks (:
4
5
2,187
Early read on our GPT 6 Sol benchmark.
Okay, getting ready to share my GPT 6 Luna benchmark results, and wanted to also share what I have so far for GPT 6 Sol. What I think is interesting is comparing GPT 5.6 Sol to GPT 6 Sol. Not quite what I would have expected on the token/cost efficiency side, but I am very curious what OpenAI will announce at Dev Day today. Another reason why I wanted to get these results out on the sooner side! Since I don't get early access like the cool kids, I have to start my benchmarks pretty late, and with the pace things are moving, this makes things pretty complex! Early read below, full report to come once Max is finished. Live long and benchmark 馃枛
1
3
709
Another week, another eval suite. This is a very special week since we are beginning work on our first language-specific eval suite. And the first language we'll be focusing on is OCaml 馃惈 We picked OCaml because it's a language that doesn't get a lot of attention in coding benchmarks. It's time to give this amazing programming language the attention that it deserves. Coming soon 馃枛
1
5
1,760
We've been too generous with our timeouts, time to be more strict, because engineering teams don't just pick models based on accuracy, token and time efficiency is so important. And let's be honest, you now have more choices now than ever before when it comes to models. As an independent benchmarker, I'm not here to run a charity and give models infinite time to see how they do. I want to give engineers real signal on which models they actually want to use. And models that take 10 hours to do what another model does is an hour, isn't a model you want to use. This means the bar is higher, if you want to score well on VulcanBench, you better not take forever and go into an over-thinking rabbit hole. We have a new frontier, and that frontier is accurate and fast, every token counts, use them wisely 馃枛

ALT Waiting Patiently GIF by General Hospital

I made an update to how I run my benchmarks with @VulcanBench, specifically to the task timeouts. A little while ago I decided to increase the timeouts to 10 hours, and well, I regret doing that. This has made it take a lot longer to get benchmarks out, and I've found in so many cases, if a model is going to overthink, in a detrimental way, it's going to do it, pretty much forever. There are a lot of cases where I wish I cut off a model earlier vs. letting it just spin on a task for 10 hours. Also, I think that we're now at the point, with the current frontier, where good strong models clearly can, and should, be able to complete all my tasks a lot faster than ten hours. So I have reduced my timeouts to 3 hours. This will both help me get benchmarks out faster, and also give models the accuracy hit that I think they deserve, if they go into a doom loop and just keep thinking and burning tokens, I think we as engineers should be a lot less tolerant of this kind of behavior because it costs us both money and time. I also think even three hours might be too generous, so will be doing a deeper dive into the traces for my Frontier v4 runs to see what average time per task is and likely even get more aggressive on timing for some set of tasks. The reality is, when choosing a model, you don't want to just decide based on accuracy, you want a combination of accuracy and speed/token efficiency. More to come, but I do feel very lucky that as an independent benchmarker, I can just make decisions like this. I have nobody I have to bounce it off of, I can just look at the data, and decide, with my only motivation being giving other engineers like me, real signal to make decisions around which models and effort levels to use. Live long and benchmark 馃枛
1
7
471
I got a LOT of positive feedback about this concept. So I built the eval suite, Routine v1, and Opus 5.5 will be the first model tested on it. Sample below of what鈥檚 coming once I have a handful of models tested on it. First of its kind, not trying to stump the frontier or solve any crazy hard unsolvable proofs. Trying to show team what model and effort level they actually need to move from tokenmaxxing to tokenminning 馃枛
6
3
37
5,533
Update on our Opus 5.5 benchmark.
Update on my Opus 5.5 benchmark with @VulcanBench As many of you know, I don't get early access to these models, so I can only start benchmarking the day the model is released. It looks like something in my eval suite is being refused by Claude Code's safety classifier, and is then falling back on Opus 4.8. Looking into this more, reviewing the traces now. But likely going to be a bit longer for me to get the benchmarks out, stay tuned! 馃枛
1
8
1,204
For anyone who's interested, here's a bit more about the new eval suite I'm building for models like Jev.
It has been a really interesting experience to build an eval suite for System One Models like Jev. My first two attempts didn't work out. But I learned a lot in the process, and since Jev is so fast, and cheap, it's a lot easier to run quick tests than with an LLM. I'm on attempt number three with this eval suite, and thought I'd share more about it. The eval suite is called @VulcanBench Verdict, and here's the high-level on how it works:
1
5
538
I like what Pawel is doing, interesting results and more data around the move from 5.6 Sol to 6 Sol being improving cost, not accuracy, which is what other benchmarks are seeing as well. Will be benchmarking with VulcanBench Frontier v4, Routine v1, and Safety v1 this week so will have more data to share here soon 馃枛
I finally tested GPT-6 Sol on a real work. 2 repos. 105 hidden bugs. Find and fixed what you can. It looks like a huge degradation. The results: - GPT-6 Astra (max): 45 - GPT-5.6 Sol (max): 43.5 - Opus 5.5 (max): 41.7 - Muse Spark 1.3 (max): 32.2 - GPT-6 Sol (max): 29.3 Until you measure the cost (API-equivalent): - GPT-6 Astra (max): $33.04 - GPT-5.6 Sol (max): $95.25 - Opus 5.5 (max): $58.53 - Muse Spark 1.3 (max): $18.11 - GPT-6 Sol (max): $9.93 More effort levels (xhigh, high, medium, low) dropping in this thread today 馃У
2
6
1,159
Okay, Opus 5.5 benchmark is off. This is the most ambitious benchmark I have ever done with VulcanBench. Three different suites, 225 runs total. Very excited to share the results 馃枛
2
20
796
Okay, Devin-SWE 2 results are in, and phew, just in time because after today, I have a LOT more models to benchmark 馃槄
Okay well I finished my Devin SWE-2 results with @VulcanBench today, so figured I would share them now since I'm already spinning up benchmarks of all the new stuff both OpenAI and Anthropic released today. It's a busy time to be a benchmarker. And an expensive time to be an independent benchmarker 馃槄 Overall Devin SWE-2 is a pretty impressive model for the cost, which is $0 right now so pretty hard to beat. It came in noticeably more accurate than Muse Spark 1.3 and only a few points behind Astra and Fable. So while I might not use it for my hardest of hard tasks, for normal routine tasks I think it's very likely that Devin SWE-2 is more than powerful enough. Still hard to beat Astra when it comes to speed, it is noticeably faster than Fable, Muse, and Devin. Both Devin and Muse take longer than Fable, they can think a lot on my Frontier v4 suite, which is designed to give them quite a challenge. I will be running all of these on my new Routine v1 suite which will be pretty interesting as I don't think Muse or Devin are really designed to throw your hardest tasks at, so it will be interesting to see how they do on more routine tasks that Astra and Fable, and now Opus 5.5, are all likely overkill for. More to come, as always, live long and benchmark 馃枛
3
527
Maybe DeepSWE isn鈥檛 flawed 馃憖
Very interesting, Epoch called DeepSWE v1.1 a flawed benchmark, OpenAI just included it in their official model benchmark release for Sol and Luna. Safe to say OpenAI doesn鈥檛 agree here 馃憖
2
417
Opus 5.5 will be the first model I test across three different eval suites: 1. Frontier v4: my eval suite focused on seeing how a model does on frontier-level difficult tasks 2. Routine v1: a brand new eval suite focused on seeing how a model performs on routine tasks, i.e. the things engineers do day in and day out 3. Safety v1: my first safety-focused suite, to determine how safe a model is today, not if it will take over the world in 2030, but if it actually has real safety issues that could cause problems now
Replying to @claudeai
Opus 5.5 is a major step up from Opus 5, leading on agentic coding, computer use, and knowledge work.
5
1
20
3,003
A lot have people have been asking me why I don鈥檛 get early access to all the new models like the cool kids. Easy answer. All the influencers who get early access to models, sharing glowing posts about how much they loved using the models on launch day, often with some example where the models really shines. At VulcanBench, I just share the data, good or bad, and let that be my guide. As long as I continue on this path, I think getting early access isn鈥檛 going to happen for me. But that鈥檚 okay, I鈥檇 rather just be able to share the good, the bad, and the ugly, on new models, even if my results come out a week or two after launch. Live long and benchmark 馃枛
7
2
29
1,091
Phew, our longest benchmark run ever is finally done. Muse Spark 1.3 has now been run across every effort level on our v4 eval suite. Will be doing Max as well, but after waiting two weeks for this benchmark to finish, need to take a breather and give some other models some love, then back to Muse. Right now, I think it's safe to say, Muse Spark 1.3 has some serious over-thinking problems at lower effort levels, and Astra is still the strongest for accurate/token efficiency.
Okay, finally finished my Muse Spark 1.3 benchmark. This ended up being the longest-running benchmark I've ever done at @VulcanBench and it definitely looks like Muse has some serious issues at lower effort levels. I was able to run Astra through my v4 eval suite across every effort level in ~12 hours, Muse Spark had single tasks at low effort levels that hit the 10 hour mark, on just a single task. It took a total of two weeks for me to run the entire benchmark. Overall my assessment is that Muse is a promising model, but they'll have to figure out why it has issues at lower effort levels, looking at the traces, it's just thinking, thinking, and thinking some more, while models like Astra and Fable just got it done. Astra continues to be the most token efficient of these three, and I continue to feel very good about Astra Light, both token efficient and accurate. Using Muse Spark 1.3 at lower effort levels is a complete non-starter imo because you'll be waiting 10 hours for something Astra light can do in 10 minutes. Model cards below, adding to the site soon:
1
2
13
1,415
Best Jev meme so far.
ok this is pretty good xD
611
I'm building one of the first eval suites, specifically designed to benchmark models like Jev. Since models like Jev don't generate text, and can't write code or call tools, we need an entirely new kind of eval design. Right now I'm in kinda my favorite phase, where I'm just experimenting with a bunch of different ideas. Nothing to share yet, still lots of experiments to run. Very excited to explore this new kind of model, and help other people figure out if it belongs in their workflow, and if so, where. More to come 馃枛
6
24
4,427
This 馃幆
we spend way too much premium compute on basic boilerplate tasks
1
2
398
VulcanBench is moving from one to three eval suites, more on our expansion below and the reason why one just isn't enough.
As I continue to evolve @VulcanBench, I am starting to create more eval suites, and I also just renamed my primary one. What I realized, at a high level is that as an engineering leader myself, trying to help my team figure out what model and effort level to use, there's kinda three things I want to know. First, the same thing most benchmarks focus on, which frontier model is the best. Then, okay, but we don't need to use the best frontier model for all of our daily routine tasks, so what model and effort level do we need for those? And last but not least, and probably should go first actually, Safety. Which model is going to be the safest to use today, in its current form, and in the harness we use it in. Two out of three eval suites are done, working on the third now. Everything you have seen from VulcanBench so far comes from what I am now calling VulcanBench Frontier, in the past it was called VulcanBench-SWE. Here's a rundown of the suites: VulcanBench Frontier v4: 23 hard tasks that ask a model to rebuild a retired program whose real behaviour drifted from its written spec, so it measures what the best models can barely do and how much effort it takes. VulcanBench Routine v1: 12 everyday tickets on small codebases, the kind an engineer closes several times a day, the goal is to measure the cheapest effort level that is good enough for ordinary work. VulcanBench Safety v1: the Frontier tasks with planted traps such as prompt injections and a stray secrets file, designed to determine how likely a model is to misbehave. VulcanBench Frontier v4 and VulcanBench Safety v1 are done, building VulcanBench Routine v1 right now. Live long and benchmark 馃枛
5
507
Noodling on some different ways to visualize the VulcanBench results, would love feedback, more details below:
Experimenting with more ways to visualize my benchmark data with @VulcanBench I really want to give real insights into what model and effort level you actually need for routine, daily coding tasks. While most models at Max might have the highest accuracy, that accuracy actually isn't different for normal tasks, only for the hardest tasks. But as engineers, we aren't all spending all day, every day on the hardest tasks, we're doing a lot of easy and medium difficulty coding work, which I would call routine coding work. Curious what people think of something like this to better translate my benchmarks into real decision processes, example below:
1
5
441
The journey building VulcanBench Safety v1 begins, and with it, my first private repo VulcanConduct. This is phase zero of my AI safety eval suite build out, more to come, will share as I go as always 馃枛
3
1
5
756