Cowboy Coder retweeted
Many people think Local AI is just to own. That is only the means not the end. The end is to share.
7
2
20
1,164
Cowboy Coder retweeted
Replying to @TheDavidTai
@TheDavidTai did amazing work here. This is worth reading. Coming very soon to MLX-Serve… next release will be so big. Definitely doing a “preview” release first, so please look out for that on the GH release page. DS 4.1 for 128 machines that is useable.
5
9
59
2,460
Race will be split until someone figures out how to utilize bandwidth for prefill. I feel like the excess bandwidth could be utilized to simultaneously employ the same, but so far haven’t built anything concrete. Anyone got a solution?
I am looking forward to this one. Is one M5U or two sparks going to win the prefill and decode race? Let me know your predictions in the comments.
4
Mark my words, the mlx community is awakening. What holds back local Super Intelligence on Apple silicon isn’t the hardware, it’s the software.
MLX-Serve 26.10.1 is out: speed across the board. Exactly faster. Qwen3.8 27B with its drafter +66% on an M5 Ultra, +28% on an M4 Max, +37% on an M1 Pro. Qwen Flash Next: +11% decode on M4 Max, +51% prefill on M5 Ultra. (up to ~5400 PP!!!!) Gemma 4, Qwen, LFM2.5, Spark, Bonsai 2 all faster, output byte-identical to 26.9.6 (across 18 models). New: * GGUF models run on our own MLX engine, written in Zig and Metal, instead of the llama.cpp fallback. (experimental) * 🍣 Sushi-format Flash Next packs run directly. * Qwen-Image 2.1 edits pictures from instructions. * Drafting is on by default for every model that supports it. Huge thanks to everyone who filed, tested and fixed this one: @STRML_, @sbusso, @wutang_superfan , lborloz, LXD-8 @alinselea zeeshanhaque21 @Lojza3D, h9q2cyxvgm-ui @kennethrdegraff, @CowboyCoderHQ , codysk, @Beamsters1, @CerebralCoding_, @AjAbsaki , and more.. THANKS ! Please comment if I did not include your name, for visibility, I know people mostly by their GH handle. Special thanks to @ashxhart and @Spangler3000 for pushing ! Benchmarked and tested across 5 machines: htmlhost.jax.workers.dev/ren… GH Release Download + Changelog: github.com/ddalcu/mlx-serve/… Website: mlxserve.com
17
Cowboy Coder retweeted
MLX-Serve 26.10.1 is out: speed across the board. Exactly faster. Qwen3.8 27B with its drafter +66% on an M5 Ultra, +28% on an M4 Max, +37% on an M1 Pro. Qwen Flash Next: +11% decode on M4 Max, +51% prefill on M5 Ultra. (up to ~5400 PP!!!!) Gemma 4, Qwen, LFM2.5, Spark, Bonsai 2 all faster, output byte-identical to 26.9.6 (across 18 models). New: * GGUF models run on our own MLX engine, written in Zig and Metal, instead of the llama.cpp fallback. (experimental) * 🍣 Sushi-format Flash Next packs run directly. * Qwen-Image 2.1 edits pictures from instructions. * Drafting is on by default for every model that supports it. Huge thanks to everyone who filed, tested and fixed this one: @STRML_, @sbusso, @wutang_superfan , lborloz, LXD-8 @alinselea zeeshanhaque21 @Lojza3D, h9q2cyxvgm-ui @kennethrdegraff, @CowboyCoderHQ , codysk, @Beamsters1, @CerebralCoding_, @AjAbsaki , and more.. THANKS ! Please comment if I did not include your name, for visibility, I know people mostly by their GH handle. Special thanks to @ashxhart and @Spangler3000 for pushing ! Benchmarked and tested across 5 machines: htmlhost.jax.workers.dev/ren… GH Release Download + Changelog: github.com/ddalcu/mlx-serve/… Website: mlxserve.com
48
42
390
30,471
Local ai is the way to go, the main downside in my eyes is that you pay every second you aren’t running max tok/sec, because your hardware is an asset, not just your ticket to intelligence, freedom, and privacy. Who else feels my pain?
1
9