I made an update to how I run my benchmarks with
@VulcanBench, specifically to the task timeouts.
A little while ago I decided to increase the timeouts to 10 hours, and well, I regret doing that.
This has made it take a lot longer to get benchmarks out, and I've found in so many cases, if a model is going to overthink, in a detrimental way, it's going to do it, pretty much forever.
There are a lot of cases where I wish I cut off a model earlier vs. letting it just spin on a task for 10 hours.
Also, I think that we're now at the point, with the current frontier, where good strong models clearly can, and should, be able to complete all my tasks a lot faster than ten hours.
So I have reduced my timeouts to 3 hours. This will both help me get benchmarks out faster, and also give models the accuracy hit that I think they deserve, if they go into a doom loop and just keep thinking and burning tokens, I think we as engineers should be a lot less tolerant of this kind of behavior because it costs us both money and time.
I also think even three hours might be too generous, so will be doing a deeper dive into the traces for my Frontier v4 runs to see what average time per task is and likely even get more aggressive on timing for some set of tasks.
The reality is, when choosing a model, you don't want to just decide based on accuracy, you want a combination of accuracy and speed/token efficiency.
More to come, but I do feel very lucky that as an independent benchmarker, I can just make decisions like this. I have nobody I have to bounce it off of, I can just look at the data, and decide, with my only motivation being giving other engineers like me, real signal to make decisions around which models and effort levels to use.
Live long and benchmark 馃枛