DeepSeek V4.1 Flash Is the New King of Local Models.
After reading this thread, you’ll understand why Dario and Sam want to ban free, open-source models.
I designed a set of brutally difficult benchmark tasks. Here were the contestants:
Local models running on 2× DGX Spark:
- DeepSeek V4.1 Flash
- GLM-5.3-Flash
- Qwen3.8 Flash Next
- DeepSeek V4 Flash Vision
API models:
- GPT-5.6 Sol
- GPT-5.6 Luna
- Grok 4.6
- DeepSeek V4.1 Flash
And one contestant entered purely for entertainment:
- Qwen3.8 27B Q3_K_XL running on an RTX 5070 Ti
The results would break Dario’s heart. Five reviewers scored every submission blind.
- DeepSeek V4.1 Flash via the cloud — median score: 98
- Local DeepSeek V4.1 Flash, quantized to EXL3 at 2.9 bpw — median: 96.5
- Qwen3.8 Flash Next — median: 93.5
- GLM-5.3-Flash / GPT-5.6 Sol (tied) — median: 93
Our sun god Sol couldn’t beat a quantized local model.
- GPT-5.6 Luna — median: 92.7
- Grok 4.6 — median: 92
- Qwen3.8 27B — 90
(“You’re telling me that models rivaling the frontier systems I spent hundreds of billions of dollars building are now being handed out for free on every street corner?
What happens to my $2 trillion valuation? What happens to my dream of becoming the father of all humanity?
Clearly, we have no choice. Open-source models must be banned!”)
The biggest surprise was the quantized DeepSeek V4.1 Flash. We’re talking about a bits-per-weight figure that starts with a 2, yet it stayed right behind the champion throughout the benchmark and comfortably secured second place.
It beat GPT-5.6 Sol and Grok 4.6. The quality loss from quantization was so small that it was practically negligible.
Grok 4.6 is actually an excellent model. I believe it’s stronger than Sol, but it has a familiar bad habit: it doesn’t like checking its own work. That cost it a lot of points here.
Its underlying capability is still formidable, and it’s much faster than the Codex family. Grok held steady at roughly 70 tokens per second, while Luna and Sol generally hovered around 20–30 tokens per second.
It feels as though they’re being deliberately throttled to push you toward buying API credits instead of relying on a subscription. But that’s fine—someone else is already giving Sam a headache for us.
The champion ran at around 300 tokens per second, with every one of its subagents running at the same speed.
Among the local models, Qwen and DeepSeek V4 were the fastest, sustaining roughly 45 tokens per second. DeepSeek V4.1 maintained around 39 tokens per second during long, single-stream sessions.
That makes V4.1 the likely choice for my main workhorse going forward.
GLM is harder to recommend on speed: it generally stayed below 20 tokens per second.
My recommendation:
DeepSeek V4.1 Flash > Qwen3.8 Flash Next
Those are the only two models you really need to download.
Installing V4.1 effectively gets you V4.0 as well. It rarely refuses requests, so in most cases, you won’t even need a separate uncensored model.