Insane developments have been happening on Qwen 3.8 Flash. In 2 weeks, suddenly everyones computer can run Qwen 3.8 Flash.
I wanted to find out what happened and why everyone in local AI is using this model now.
Every number below is a receipt I verified, my own hardware or community, and the ones from my fleet live in my bench repo. Here is what I found.
WHY THIS MODEL
Qwen3.8-Flash-Next is a 180B parameter MoE that only activates around 6B per token.
On Artificial Analysis it scores 40 on the intelligence index against Claude Opus 4.8 at 42.
It beats Opus on Terminal-Bench 4.0 (25 percent against 22), AutomationBench (56 against 46), and GDPval. Opus still wins Humanity's Last Exam (49 against 38) and SWE-bench Pro (69.2 against 62.5, each vendor's own harness).
My read: Opus-class on agentic and coding work, one step behind on the hardest reasoning.
People call it Qwen 3.8 Flash. Same model.
WHY IT RUNS ON ALMOST ANYTHING EXPLAINED
Traditional LLM understanding is everything has to run on vram alone. Model doesn't fit VRAM - you can't run it.
The Chinese AI companies broke the conventional datacenter model basically, by showing people off-loading critical parts of the model out of it is possible, but there was a catch -
Everyone knows that RAM offload is slow, and it was true. A dense 70B at 4-bit is around 35 GB of weights and every token touches all of it.
From system RAM at 60 GB/s that is about 1.7 tok/s.
Flash-Next breaks the rule because of its shape. 512 experts per layer. The router picks 10, plus one always-on shared expert.
Per token you read maybe 2 percent of the expert bank, around 2.2 GB at 3-bit. Same RAM, same bandwidth: 60 divided by 2.2 is around 27 tok/s as a pure-RAM streaming bound. Result: Laptop runs suddenly become possible.
Sparse active set + Treat your whole computer as RAM.
The new engines stack three things on top:
The cache hierarchy. Expert usage is lopsided. Some fire constantly, most rarely. The engine treats your PC like a CPU cache: hot experts pinned in VRAM, warm ones in RAM where the CPU computes them, cold ones streamed from SSD. Speed becomes hit rate, not bandwidth.
The n-gram table. 51B of the 180B parameters are a lookup table. Not matrix math, an index. It sits on your SSD and streams a few rows at a time.
MTP speculative decoding. A draft head guesses, the big model verifies several tokens per pass, and corrrect guess is a free speed up.
STRATA CHANGED THE GAME (AGAIN)
The engine that went viral is Strata. MIT-licensed, built for this one model family, one-click install. It serves an Anthropic-compatible API on localhost, so Claude Code and Codex CLI point straight at it. That is the real adoption driver. Same 5090, llama.cpp: 15 tok/s. Strata: 160 tok/s. Engine, not model.
The quant side moved too. ISTA-DASLab's GSQ-RCO assigns bits per tensor by measured sensitivity. Their 3.5-bit build (83.6 GB) scores near or above the BF16 base on most evals. A quarter of the size, matching quality.
THE FULL MAP
8 GB VRAM
RTX 4060 laptop (Strata, Q3): 30 tok/s at 192k context, 130W from the wall. Tono_Ken3's receipt. A gaming laptop is now an inference server.
GTX 1080: Apparently this runs at 20 tok/s painfully, but it works lol
12 GB VRAM
RTX 5070 + 64 GB RAM (Strata, Q2_0): 94 tok/s decode, 2,650 tok/s prefill at 32k. The project's reference rig.
24 GB VRAM
RTX 3090 + 64 GB RAM: Two lanes. llama.cpp UD-Q2_K_XL with expert offload: 28 tok/s at 65k context, sustained over 1.59M production tokens (eirrann_art). EXL3 2.5bpw with MTP: 38.6 tok/s, full 262k native context, needs around 59 GB host RAM (r0b0tlab).
RTX 4090 + 64 GB RAM (Strata, IQ3_S): 60 tok/s decode, 2,500 tok/s prefill. 5000e12's receipt.
RTX 5090 (Strata, IQ3_S): 160 tok/s, up from 15 on llama.cpp. EpicMaan's receipt.
32 GB VRAM (LEGACY)
2x V100 32 GB (64 GB total): 60 to 88 tok/s at 262k context on pentacoxian-dev's custom quant. ryu15's receipt. The e-waste lane is real.
128 GB UNIFIED (DGX SPARK)
1x Spark (my grid, EXL3 native): 107 tok/s count, 71.7 prose, 15 of 15 needle at 243k context.
2x Spark TP2 (my grid): 68 tok/s count on stock NVFP4. 138 on hibrid48's stock-incompatible fork. Same chips, the recipe alone is worth 2x.
64 GB+ UNIFIED (MAC)
Macs: two engines carry it, not just the GGUF lane. TensorFold runs the full 180B on M1 through M4 using oQ checkpoints, with the n-gram table scaled and dense projections moved to the matrix units before M5.
oMLX, which owns the oQ format, serves it too, with SSD-paged KV caching and a menu-bar app. Floor for the 180B is still 64 GB unified.
Under that, the Mac lane is the Qwen3.8-27B dense: 131-160 tok/s on an M5 Ultra, 16 streams at 32k on a 64 GB machine.
THE 32 GB QUESTION
Can it run on 32 GB of RAM? Two weeks ago the answer was no, the floor was 48 GB.
Strata's spec floor is now 32 GB RAM plus 12 GB VRAM, but every receipt I could verify sits at 64 GB. The floor claim moved to 32 this month. The receipts live at 64. If you are on 32, the honest answer today is the Qwen3.8-27B dense.
This model will become a mainstay for local AI for many months to come. It's reasoning capabilities are very good in my own testing, it's not GLM 5.3 or Deepseek V4.1 level, but for a driver for most daily use.
Opus class model intelligence has arrived to almost every computer, and yes Qwen 3.8 Flash is stronger than Qwen 3.8 27B dense in my testing.
Sources, receipts, and how tos in the replies 👇
The DGX Spark has gone up in price, and that is the bad news. At a new price of $7000 , alot of people are going to ask - can you do much with just 1 DGX Spark?
The good news is you can run more frontier-class stuff on it than ever.
60-70 tok/s (322 in concurrency) on Qwen 3.8 Flash is what you can do on it. Yup, you read that right.
I spent last night testing the newest proof:
@vr8vr8 's single-Spark recipe for Qwen3.8-Flash-Next.
Full agent grid, every lane.
I was lucky enough to be informed of the recipe early. This is v5.1, and it is a preview.
The recipe will get better, bear that in mind.
HOW DOES HIS RECIPE ACHIEVE SUCH SPEEDS
Qwen3.8-Flash-Next mixes regular attention layers with recurrent state layers. That hybrid architecture makes speculative decoding expensive.
A normal draft means saving and restoring the recurrent state for every guessed branch. Memory and time per guess. Guess wrong and all of it was wasted.
His v5.1 ports RecoverSSM. Instead of forking the recurrent state per draft, the engine saves one checkpoint, runs all drafts against it, then replays only what got accepted. Deeper drafts stop costing memory.
That single change is why the KV pool grows to 876k tokens and the seat count doubles from 8 to 16 on the same silicon.
MY GRID
Same frozen clocks as every grid this desk runs. Empty context, decode after first token:
Lane1 stream2 streams4 streams
Count to 20081.7146.2244.4
Hash map explainer58.289.6134.3
50 Python clamps83.9132.6236.6
JSON GPU stats69.3111.9207.1
Context window, 262k native. Decode after packed filler: 48.6 at 8k, 48.4 at 32k, 50.2 at 131k, 49.4 at 262k. Flat, wall to wall. The box holds its entire window without sagging. Needle retrieval, three positions at four depths: 12 of 12 found, deepest rung 255,787 actual tokens.
Agent lanes at 35k context, because tok/s on an empty cache is not the agent experience. A real build turn writes 2,048 tokens of a single-file HTML game at 77.2 tok/s after a clean tool call. A research-then-build turn lands at 57.1. Tool calls parse exactly, no XML drool.
Concurrency ladder, prose clock: 50.7 at one, 122.9 at four, 189.6 at eight, 282.6 at sixteen. His table claims 322 at 16. Mine reads lower. Both numbers are provisional, and even at mine that is sixteen concurrent agents on one desk-side box at 0.20 joules per token.
ONE DGX SPARK VS TWO DGX SPARK
It's been more and more impressive what runs on a single DGX Spark, after trying recipes from vr8vr8 and
@ViC305 and
@mr_r0b0t . Please please check out those accounts if you have or are looking at a single DGX Spark setup.
To be clear two sparks are still the sweet spot, and gives considerable improvement, but with the price increase - the dual-spark crosses the psychological $10k and maybe a bit uncomfortable.
I have my dual-Spark grids on the same weight family, same clocks, so the trade-off can be clearly mapped.
Single stream, the single box reads 59 to 71 percent of the dual depending on lane.
My dual grid ran his v4 recipe, this is his v5.1, so some of the closeness is the engine improving.
The dual also still owns depth. Its YaRN stretch holds 1M context (though I added that myself). The single box is 262k native. I have not tested YaRN on this recipe yet.
But my dual serve configured at 8 streams max. The single box recipe does 16. The KV pool size is still much larger on TP=2 obviously. (more context per stream)
A PERSONAL NOTE
The recipe ships an optional patch called hermes-chat. I helped with the patch 😃 My dual grid on his earlier recipe hit harness friction where thinking defaulted on when my agent harness omitted a field, so I contributed the fix upstream and he shipped it in this kit with credit. Ty!
Second time I have shipped my own patch inside someone else's recipe. It works both directions, and the tool-call parsing that used to leak XML into content is clean now too.
WILL IT STACK WITH TENSORFOLD?
The speed in this recipe comes from RecoverSSM: a smarter memory layout that lets vLLM run deeper drafts without burning KV per branch. Mia's TensorFold recipe for the same model on the same box gets its speed from the engine: tree-verified speculative decoding that commits more tokens per verify round than vLLM does. Her single-stream prose lands at 62.4 tok/s against my 50.7 here.
Those are two independent wins. One is where the memory goes. The other is how the verifier works. Nobody has tried both at once on this model.
RecoverSSM already landed in upstream vLLM as a PR. TensorFold is MIT and speaks the same weight formats. The technical barrier to porting RecoverSSM's checkpoint-and-replay into TensorFold's verify loop is not zero, but it is not a rewrite either.
If they stack even partially, you get the 16-seat concurrency of this recipe at the single-stream speed of Mia's. On one $7,000 box.
That is the real question this grid left me with. The model is done improving until the next checkpoint ships. The infrastructure around it is not.
Sources, receipts, and my grid benchmark in the replies 👇