The Qwen 3.8 Flash Next speculative decode shootout!
We can get speeds consistently above 80-90tps for contexts up to 128k, hitting over 107tps on realistic coding workloads by make more permissive speculative decoders.
Here I'm comparing three different types of speculative drafters and how they impact accuracy and decode tokens per second. Speculative drafting like MTP guesses extra tokens better drafting is the biggest knob we have at increasing tokens per second. You can either build better drafters or you can build more permissive verifiers for the drafters.
First we have the previously implemented typical verification (use good enough tokens), now we are adding cascade verification (how close is the drafted token to the verification's desired token) using a straight forward version (OTP) and a version that corrects for errors (TokenV3).
On the first graph which shows decode tps (prefill is unchanged so uninteresting), you can see first that OTP is very fast but you're sacrificing massive amounts of accuracy. We don't really want to do that, we want to pick the two that preserve accuracy.
On the second graph which shows HumanEval+ accuracy (left bar) and task completion* (right bar), you can see here that for our tests, TokenV3 with the 0.95 setting and Typical Verification with the 0.2 setting bot have a ~20+% increase over exact verification with TokenV3 slightly being faster than Typical 0.2.
So you can see that we can get increased performance with less exact verifiers without losing much accuracy.
*task completion is the percent of tasks didn't infinite loop
github.com/youssofal/MTPLX/p…
github.com/youssofal/MTPLX/p…
Sep 10, 2026 · 12:34 PM UTC
3
2
27
2,968
Typical Acceptance Paper
arxiv.org/abs/2401.10774
1
3
123
Spec Cascades Paper
arxiv.org/abs/2405.19261
1
2
119




