The latest improvements on vLLM enhances serving @Alibaba_Qwen Qwen3.8 Flash Next. Improvements are coming to the recipe!
762 commits. 315 contributors. 104 first-timers. vLLM v0.30.0 is live. 🎉 Highlights: 🤖 Hybrid-attention hot paths: Kimi K3 streamlines KDA, AttnRes, and MLA; DeepSeek-V4.1-Flash adds MXFP8 KV and async Engram; Qwen3.8-Flash-Next fuses QSA/PLE and cuts sparse-GQA overhead 🗄️ HiSparse adds a host tier beneath sparse-MLA decode; under GPU pressure, only top-k misses return to a per-request hot buffer 🛠️ Model Runner V2 brings EAGLE3-style drafts to pipeline parallelism and extends adaptive verification to every draft-model speculator through online acceptance estimation (#50514, #52228) 🖋️ Dual-key Gumbel-max watermark generation and detection, with per-request opt-out and speculative-decoding support 🆕 New models include GLM-5.3-Flash, K2-Horizon, Cohere Compass, and Bailing V3 VL ⚡ Fast Start keeps post-quantized, TP-sharded weights in a per-GPU daemon; restarts map them over CUDA IPC with --load-format ipc_cache Thread 👇
Made with AI

Sep 25, 2026 · 7:32 AM UTC

8
6
65
26,722
Sort replies: Relevant Recent Liked
I managed to get 86 max :) what’s yours? Dropping my recipe soon too 😂 vllm 0.30.0 rocks
1
2
490
Oh that's good! I'm still working on it
1
1
363
Serving patch beats a new name.
193
Need recipes for ThinkingCap and Swift finetunes for Qwen Next. Overthinking is still a problem for me.
133
This is great news. Been waiting for that sweet, sweet, engine upgrade. Hopefully I'll be able to get storage tier KV caching working :) thanks a lot @jvr0x!!
157