Comparing MiMo-V2.6-Flash in DS4 to the Xiaomi API. More work to be done on long context :) Still cool that a single M3 Ultra can compete on decode performance. Also, prefill on M5 Ultra should be much better.
4
812
And here's MiMo-V2.6-Pro-RL in DS4 on 2x M3 Ultra
3
10
6,174
Here's MiMo-V2.6-Flash-RL in DS4 on M3 Ultra
4
2
23
2,397
Pre-orders for M5 Ultra 256GB already at 16-18 weeks shipping. Hopefully this is not bad news for the 512GB version...
2
1
16
2,356
You can run the full DeepSeek V4 Pro 0813 in native precision at a fairly decent speed, locally on 2x Mac Studios using DS4. Also worth noting that this model beats Mythos Preview at CyberGym (83.3 vs 83.1)
2
7
38
5,041
Some throughput benchmarks on one and two machines. Tensor parallelism via RDMA over Thunderbolt improves generation 1.5x and prefill 1.8x.
529
Per-token divergence from the full model, by span. It picks the same next token 95.5% of the time on tool calls and 87.4% on the final written answer, with the thinking spans at 78.1% carrying most of the cost.
2
3
781
Relative perplexity on 8 codebases across 5 languages, against the full 1.56TB release. The Linux kernel stays within 5% of the baseline.
1
3
698
@antirez @ivanfioravanti some tweaks to the DS4 TP support in case its useful github.com/kernelpool/ds4/tr…
2
2
29
2,698
Benchmarks for the Kimi K3 2bit quant thanks to @ivanfioravanti's llm_context_benchmarks tool. Performance holds up nicely over longer context thanks to KDA/MLA 🚀
Kimi K3 (2.8T params) running locally on two Mac Studios, pi agent + mlx-lm
2
4
9
3,525
Kimi K3 (2.8T params) running locally on two Mac Studios, pi agent + mlx-lm
21
13
276
60,669
Here's a mlx-lm (PR #1093) vs llama.cpp (b8660) comparison for Gemma-4-26B-A4B on M3 Ultra
Direct comparison of NVIDIA RTX 5090 to M3 Ultra 👀 Small models should always be faster on the 5090. The best perf for large models is to use both together (more on that soon). Using llama.cpp isn’t super fair given the performance is not great on Apple Silicon. MLX is better
3
2
29
8,305
Plenty of great improvements in mlx-lm 0.30.6! Here’s Kimi-K2.5 running on 2 x M3 Ultra, up to 128k context 🚀
Latest mlx-lm is out: - New models: Kimi K2.5, Step3.5 flash, LongCat Flash lite thanks to @kernelpool - Support for distributed inference with mlx_lm.server thanks to @angeloskath - Much faster and more memory efficient DeepSeek v3 (and other MLA-based models)
2
10
46
8,522
Here's another one from Kimi-K2.5-3bit, running on a single M3 Ultra. I was only able to test up to 8k context without MLA absorption.
1
103
Here's Kimi-K2 distributed
2
2
4
1,513
It runs a little slower, but its much better than running out of memory after about 20k tokens :) However, I did something similar to DeepSeek 3.2 earlier this week, and did benchmark that:
2
4
434
Always fun to see speed runners pull off arbitrary code execution exploits in real time 😃 piped.video/watch?v=aG7_bypp…
3
15
6,937
In Apple's defense, there's a carefully articulated warning in the MLX docs ml-explore.github.io/mlx/bui…
10
1,594
MiniMax-M2 3bit DWQ vs standard 3bit quant: +4.1 points on MMLU-Pro 📈 (Took a week to benchmark 😅) 🔗 huggingface.co/catalystsec/M…
1
1
1,792
Oh wait, looks like SPTM is backported to M2/M3 in macOS 26 🎉
1
7
1,302
Faster than MLX 😮 (from M3 Ultra)
3
154
We’re ready for round 2 of @offensive_con ! Come say hi and grab a pen and stickers at our booth :)
1
17
3,196
Found some reference material for Stargate project
4
1,269
First day of #HITB2023AMS starting soon (iOS/OSX security panel discussion is tomorrow morning). Let’s go! @HITBSecConf @TrenchantARC
1
4
31
4,981
Apple reminding us that at the end of the day, it’s all about the emojis
1
8
2,991
Tried to bribe my way out with an 0-day. Didn’t go too well.
17
6,598
Checking out Safari in Lockdown Mode
4
34
256
Getting ready for a long day! Check-in starts 10:30am #driven2pwn #HITBCyberWeek
7
10
Today was a good day
5
61
LOL. The day after Trump pulls out of the Paris agreement, Australian TV stations start airing this ad again piped.video/watch?v=7UcFY5fw…
3
Checking out the Grayhash office in Seoul /cc @beist
2
Lol @Allianz can't accept online payments after hours.
2
8
18
Song of the day - Epistemology (Keep of Kalessin): piped.video/watch?v=mVuxJuO7…
2
2
Someone built a quad copter of the Defcon badge. Awesome!
3
51
22
Replying to @aloria
@aloria I'm opening this on its 20th anniversary
1
1
Found this awesome book waiting for me in the mail. Thanks @brucedang and @0xeb :)
2
7
Alex (@defendtheworld) talking about mobile messaging app security @ #codegate
4
4