MASSIVE performance update to Qwen3.8 Flash Next on a single DGX Spark ๐ฅ
- 46 tok/s for prose, single stream
- 108 tok/s for prose across 4 concurrent streams
- 2,000โ2,200 tok/s prefill across all context sizes
- Full numbers in the post below
KV cache is now 1M instead of 1.4M due to stability issues. This was necessary to avoid OOMs.
This is THE best model to run on a single DGX Spark!
Expect further improvements.
Get it here:
github.com/MiaAI-Lab/Qwen3.8โฆ