introducing qwenfast
the fastest inference for qwen3.8-27b, beating vanilla vllm, hosted on h200
i use the model's built-in mtp head for speculative decoding. large batches are bound by moving the 72 mb per sequence recurrent state, so i wrote my own triton deltanet decode kernels that run at over 80 percent of hbm bandwidth. everything runs inside cuda graphs.
qwenfast: 133 t/s || vllm: 97 t/s,
without compromising quality.
gsm8k: 0.92 || ifeval: 0.98
vllm: 0.915 gsm8k, 0.96 ifeval
try it yourself:
qwenfast-demo.vercel.app
also, no token limits, burn as many as you want..