Okay, and now the model card I've been waiting all week to build and share. This has been the question on my mind since Dev Day on Tuesday, so I'm guessing it's on other people's minds as well.
GPT-6.1 Sol vs. Astra vs. Opus 5.5, across every effort level on
@VulcanBench Frontier v4.
And before I share my own insights, one note I think more benchmarks should be sharing. What accuracy score do you really need a model/effort to hit, for your routine daily coding tasks?
So many benchmarks show performance at Max effort, but I never use Max effort, never use Extra High, sometimes use High, but mostly use Medium and Low with the current frontier because that's more than powerful enough for most routine daily coding tasks.
And, we all know you don't need to use the best frontier model with the highest accuracy score. Sure you can do that, but you'll spend a fortune, and you'll find yourself in the 2025 agentic coding pattern where you give a model a task and just sit and wait.
On VulcanBench Frontier v4, I typically say that a model/effort combo between 80 - 85 is great for easy/medium difficulty tasks, 85 - 90 is a good choice for medium/hard tasks, and above 90, well, you may never need.
For me personally, I think I need model/effort combos that score above 90 maybe 5% of the time, the sweet spot for most of my routine daily coding work is model/effort combos that score in the low 80's on Frontier v4.
Do with this info what you will, we're all learning together too so take it with a grain of salt. And now, for the benchmark I've been waiting to share, see below.
Live long and benchmark 🖖