Rough M5 Ultra 80 core GPU 256GB benches:
Qwen 3.8 Flash Next, Q4 variants, MTP, 1 stream, 32k context, averaged code/prose:
TensorFold 0.3.4.1: 152 tok/s decode, 2086 tok/s prefill
mlx-serve 26.9.6: 124 tok/s, 3863 tok/s prefill
oMLX 0.7.0rc1: 105 tok/s, 3694 tok/s prefill
10
1
50
17,019
Could only get oMLX to run Q8, so here's a Q4 vs Q8 comparison for oMLX only.
Seems to be a ~12-14% cost to moving to Q8 ad the moment.
95 tok/s Q8 vs 105 tok/s Q4 @ 32k
Sep 28, 2026 · 3:29 AM UTC
2
2
629





