what's interesting is this is probably the first benchmark I've seen that feels like it relates to the real-world and it's kinda nailed it
for example, I'd be willing to wager most people don't *really* know the difference between gpt-5.4, 5.6, or 6
including software devs into that cohort, it's surprisingly hard to confidently say a model increment has a meaningful impact by just sorta vibe testing it (sometimes you can tell sometimes you can't)
so the only real way to know if a model is "better" is by looking at benchmarks like SWE-bench, OS-World 2.0, Humanity's Last Exam, Terminal Bench etc.
but unless you're a researcher or into benchmarks you can't really grasp the leaps these models are having. getting 52% instead of 51.2% on the SWE-bench is hard to gauge:
*has the model got better?*
or
*have we got really good at hitting benchmarks?*
instead I would be much more interested if labs optimised for something like DrivingBench where it's solving a clear problem and the numbers have a much more tangible meaning
you could expand this into any domain like SurgBench (
arxiv.org/pdf/2506.07603) or ButterBench (
arxiv.org/html/2510.21860v1) where actually care less about the numbers and more about whether the model *actually* did the thing
We gave ChatGPT, Claude, and Grok control of a real Toyota Corolla 🚗, steering/gas/brakes, no human driving (just a foot over the brake).
Only one model was able to complete our entire driving course. Introducing DrivingBench. 🔥⌛️🏁