Rough M5 Ultra 80 core GPU 256GB benches: Qwen 3.8 Flash Next, Q4 variants, MTP, 1 stream, 32k context, averaged code/prose: TensorFold 0.3.4.1: 152 tok/s decode, 2086 tok/s prefill mlx-serve 26.9.6: 124 tok/s, 3863 tok/s prefill oMLX 0.7.0rc1: 105 tok/s, 3694 tok/s prefill
10
1
50
17,019
Concurrency is pretty rough at the moment, but I think multiple people are working on it actively for M5U. 8 concurrent requests, no MTP, 2k context: mlx-serve/oMLX: 223 tok/s total (28 tok/s per stream) TensorFold: Doesn't support concurrency
1
3
808
Could only get oMLX to run Q8, so here's a Q4 vs Q8 comparison for oMLX only. Seems to be a ~12-14% cost to moving to Q8 ad the moment. 95 tok/s Q8 vs 105 tok/s Q4 @ 32k

Sep 28, 2026 · 3:29 AM UTC

2
2
629
I think there's a lot of headroom to be squeezed out still, and it's obviously early in the chip's life. A rough test shows the GPU can draw ~162W. GPU power, decode / prefill: TensorFold: 97 W / 123 W mlx-serve: 87 W / 130 W oMLX: 71 W / 150 W Decode maxes out at 430-460 GB/s
3
457
Sort replies: Relevant Recent Liked
Replying to @iam_agg
Hey, thanks! mlx-serve won't run Q8?
1
29
Not sure why I said that, sorry I was a bit flustered by everything wanting different quants. I just didn’t pull a Q8 for mlx-serve (because I got a oQ8e for oMLX, and found that TensorFold only works with specific things, and completely forgot about mlx-serve)
2
69