playing with Expert expansion in MOE Qwen 3.6 35B
mmlu-pro intermediate results :
dyn 20 vs static 8 experts 0.820 vs 0.805 (the same 6bit quant gguf)
LESS tokens -9% , more latency +1%
what is :
"Drafted / verify call
0.00tok/call
0 drafted · 1,071 verifies"
in my M4Pro is always zero, but i can guarantee that old fused kernels drive Qwen 3.6 35B A3B so well that has never been so fast (without losing smartness , in comparison to unsloth 6b XL )
More than 10 tok/sec with a near SOTA model likes Deepseek v4 flash on my MacBook pro with M4 pro and only 48GB of RAM using the great DwarfStar tool from @antirez and it's SSD streaming feature!
github.com/antirez/ds4