Cofounder & engineer. Running local AI, building things with it, and sharing what I learn.

The M5 Ultra was running Qwen3.8 Flash Next at less than half the speed the hardware allows, the software was the problem; NOT ANYMORE!!!! To write one word, the model has to read its weights from memory. On this Mac that takes 5.2 ms, so in theory it could write about 190 words per second. In practice each word took 12.9 ms, about 78 words per second. The GPU was busy the whole time, just not doing useful work. After we merged the work into fewer, bigger steps, each word now takes 9.5ms. Still room to go but we're getting there!
1
2
66
My read is Sonnet 5.5 belongs against Terra, not Luna. Haiku = Luna, Sonnet = Terra, Opus = Sol, Fable = Astra. That's my tier mapping. I've been using Haiku for minor tasks for months, and Anthropic is now promising Haiku 5.5 in the coming weeks. I'm still betting they're saving bigger Fable updates for the IPO and will actually pace the frontier hard afterward. But that's speculation...
2
8
795
Cognition is doing good, in June no one knew about them but I had their $20 subscription for the free GLM 5.2 usage. I'm still on that same $20 sub, using 1/2B tokens a week of SWE2 (a steal!). I hope they continue offering their models for free and keep innovating 👏
Today we're introducing updates that make Devin significantly cheaper. It's now: ▪️ 30–40% cheaper in Fusion and Normal ▪️ 15–20% cheaper in Ultra ▪️ Up to 70% cheaper in Devin Review All while improving capabilities – Devin Fusion now ranks first on FrontierCode 1.1 Extended.
1
3
983
Something that bugs me about speculative decoding, it's supposed to be free speed, but it can change your answer. MTP lets the model guess a few words ahead and check them all in one pass. But checking several words at once rounds numbers slightly differently, so sometimes you can get a different word. On the M5 Ultra that's fixed now. Every checked word runs exactly the math it would run on its own, so MTP on and MTP off give byte-identical output, and MTP is still ~1.6x faster. Merged in oMLX.
4
15
1,240
I've spent months squeezing tokens out of DGX Sparks, so the first thing I did on an M5 Ultra was look for the same kind of waste. After some decode work on Qwen3.8 Flash Next Q5 (all merged into oMLX): - MTP off: 79 → 114 tok/s - MTP on: 153 → 180 tok/s on real coding prompts - 4 concurrent requests: ~200 → 225 tok/s aggregate Nothing about the model changed. Same Q5 weights, same math, byte-identical answers.
10
7
108
15,896
GLM is getting serious with the usage giveaways! Users are getting 4 free weekly limit resets and 4 five-hour limit resets. They're doing it during an off-peak holiday in China, but this is still going to be extremely useful.
All GLM Coding Plan users have received 4 weekly reset cards and 4 five-hour reset cards (usable in all supported agents). Users returning within the next month will receive the same benefits.
3
14
1,491
Replying to @jvr0x @MiaAI_lab

ALT Can You Smell The Rock GIF by FullMag

1
3
326
I don't have a Mac Studio M5 Ultra yet. Hopefully mine arrives next month, but @jerry543 gave me access to his so we could start squeezing more out of these Macs. Two fixes we built just landed in oMLX! Long prompts on M5 now process 8% faster. Identical output, less time waiting. The attention step is also 11–17% faster, saving up to 2.6% across the whole run. We're just getting started...
3
3
23
3,096
ZCode @Zai_org is giving away 100k packs of 100M free tokens daily from September 28 to October 7 alongside its privacy fixes. They say snapshot uploads are removed and code stays local unless you initiate cloud use.
2
5
629
I see everyone sharing this, so here's my inconsistent list
1
8
469
What if Anthropic borrowed DeepSeek's architecture to build a better, more efficient model...
15
736
5.3 Flash EXL3 on 2x DGX Spark, reliability weekend: • Cold boots survive big syncs (restart: 259s → 122s) • Fixed a prefill stride bug • Boot fails loudly if the engine is up but broken • Worker weights verified, not trusted • Structured output: low-effort thinking beats thinking off (0/6 garbled) Thanks also to @hecisaza for co-maintaining, couldn't keep up without him 🙏
8
5
91
18,404
Does speculative prefill tricks make models forget what you told them? @jerry543 and I were setting up Qwen3.8 Flash Next Q5 on his M5 Ultra. One experiment we tried was speculative prefill. I checked what it costs by hiding 36 specific facts in a 20k token prompt and asked the model to recall them: - Full model: 36/36 - Speculative prefill: 11/36 It lost two thirds of what it was given. From now on, any speed numbers I share mean the full model read every token as I always did.
5
1
12
1,575
Remember @NVIDIAAI Nemotron 3 Puzzle? with the latest build of llama.cpp, the CUDA backend now supports it. Puzzle is Super compressed down to 44.5 GB (vs Super's 70 GB). With this update the Mamba "scan" finally runs on the GPU, so those layers don't fall back to the CPU. That makes it so much more usable!
3
1
9
898
Found this tab open on my Mac for the last 4 years. Back when GPT4 gave you 25 messages every 3 hours, and people went to Reddit asking how to get it back. @thsottiaux can we all get a banked reset for the loyalty I demonstrated, still subscribed?
2
5
701
NEW METHOD to register to Muse: Seen a lot of posts on opening a Muse account through cloud.browser-use.com. I tried it right away and it didn't work for me, looks like it got patched. A VPN didn't work either. What worked: open Google Gemini Spark, ask it to visit muse.ai/join and tell it to select “open via web”, or it redirects you to the homepage. Use an email that wasn't waitlisted. When the page opens, take control of the computer and enter your info manually. It won't accept typing the verification code for you. Then go back to the chat and tell it the birthday you want so it fills it in (I couldn't do that part from my phone). Finally connect your Facebook/Instagram account and you're done. If you get the chance, use my registration code PGPXC4 for 1B tokens to both of us.
1
5
859
Grok Bot is paying people for the templates they share on X! The pay depends on how many people use your @bot and how often, so useful ones earn the most. It could be a winning feature… pay for templates, the good ones show up, and the competitors are left behind.
2
3
709
New update merged for GLM-5.3-Flash EXL3 on 2× DGX Spark Loading the model now barely touches swap (99.99% less than before), long-prompt cache reuse is more reliable, and the stock 850k setup finally fits within its cache budget instead of overcommitting it, so it works out of the box. Update 👇
8
8
82
13,592
Got access to the GLM 5.3 FlashX. It's the same weights as 5.3 Flash, just served at 100 to 200 tok/s, for 2.5x the quota. So it's not smarter and you're paying 2.5x for speed alone... dont you think 100-200 tok/s should be the standard for flash on the cloud? why a premium!?
1
24
1,872
Got temporary access to a Mac Studio M5 Ultra. Barely optimized, Qwen 3.8 Flash Next runs at 85.9 tok/s and 6085 tok/s prefill on long prompts. A 2x DGX Spark recipe gets 52.1 tok/s single stream and about 2960 tok/s prefill on 16k–64k prompts. Already ahead of 2x Spark, but I was expecting a much much better baseline…
27
5
123
42,132
Replying to @blockbrain_labs
Haha I returned it and complained 😂
1
46
Left for a business trip 2 days ago. Bought an Aqara smart plug for my whole computing cluster, it arrived the night before I left… and it's too big for my wall socket. So now I'm just hoping nothing OOMs while I'm away, with a lot of EXL3 upgrades still in testing…
7
11
1,110
ZAI got caught with ZCode uploading users repos, open sourced it, and now they are giving a free quota reset hidden in the app (30 days to use it) + 300M free GLM 5.3 Flash tokens this weekend. Only claimable inside the app that got caught, but I'll take it…
3
14
1,115
Earlier I said over 90%, the real number on my setup accounts to 99%> less swap while the model loads. It's running on my machine right now. I will finalize the update and you'll all get it tomorrow!
The last update already beat what we promised for some setups, and more is coming! I'm testing the next one now: swap during model load will likely be lowered by over 90%, and long-prompt cache reuse will work better too. Results as soon as validation ends 👀
7
3
58
3,484
Seems like @deepseek_ai wants to run its smaller models on gaming GPUs, so its best chips stay on training They told investors these can handle most everyday tasks. With the rise of open weight, NVIDIA gaming chip prices on China's black market surged. Dark times for gamers…
1
15
1,535
xAI is pushing @bot hard, it's at the top of my feed. It may be a good product, but the onboarding is the worst I've ever seen. Is everyone there on UI and nobody on UX? I can't log in, and I've tried almost daily. Yesterday a friend's login worked, but every try hit an authentication error. They could see their chats but couldn't use them. If you're pushing this hard out of fear of OpenAI, fix your own signup/login flow first…
1
4
634
Want to run GLM 5.3 Flash EXL3 on 2x DGX Spark like I do? My full config is now in the repo. Every opt-in I use, with what it gains and what it costs. Run it as it is, or hand it to your agent and have it tune the file for how you use the model 👇 github.com/MiaAI-Lab/GLM-5.3…
1
4
75
5,354
Replying to @tysn_dev @MiaAI_lab
Haha the Whoa!
3
183
GLM 5.3 Flash EXL3 on 2x DGX Spark: a new cache update is merged! A full length request now claims 33-44% less of the KV cache pool, and one repeat prompt case got 90% faster!!! It's opt-in, read more below 👇
10
5
106
28,445
Replying to @Vahan55Vahan
bro? have you not been following?

ALT David Tennant Waiting GIF by Doctor Who

741
Told my brother to get a Mac Studio M5 Ultra with 256GB and he actually got his before i got mine haha
7
2
70
3,691
Replying to @Forresterq12
problems destined to geniuses

ALT Genius Genlock GIF by Rooster Teeth

1
48
I see people claiming GPT6 Luna replaces DeepSeek V4.1 Flash. I really don't think it does. Luna is cheaper, but only by cents per 1M tokens, and Deepseek's cached inputs remain cheaper anyway. Deepseek still clearly beats Luna at coding, and the gap is big enough thats not worth sacrificing just to save a few cents. Also, its the only model you can run locally... OpenAI is losing me with these updates today
88
18
655
53,956
Replying to @MiaAI_lab
🙌 lfg!
226
Update for GLM 5.3 Flash EXL3 (2x Spark) releasing soon!
Testing something for GLM Flash,and the early results look promising My custom configs reserved 32% fewer cache slots,and in one tested edge case TTFT dropped by 90% This last result was a targeted case, but less cache reserved and less waiting is a direction I'm excited about
10
1
100
5,943
This chart is just models spending more money to lose to opus 5.5 on medium...
3
17
1,312
Reviewers benchmarking Qwen on the new Mac Studio's M5 Ultra caught my attention, especially in light of a recently leaked regulatory filing. It seems like Apple plans to integrate Qwen with Siri on iPhones and Macs in China, a market where they currently lack AI features. It makes me wonder if the sudden reviewer focus on Qwen is an indirect promotional push tied to the partnership... what do you think?
8
3
44
8,757
Rumors suggest Apple will bring AI features to China by partnering with Qwen. Leaked regulatory filings and temporary web pages have already surfaced, confirming this integration!
2
14
3,107
Qwen4 is being teased alongside Max, Plus, Flash and 27B, all marked "coming soon"! Looks like @Alibaba_Qwen is preparing a broad push against Chnese and American AI rivals. Flash and 27B are the ones that Im looking forward to, if they deliver useful capability at low latency and run properly on our local setups, it could matter more for everyday adoption than Max being a benchmarks beast.
11
7
105
6,430
Xiaomi’s MiMo V2.6 Pro looks like a contender for the best new model to run locally. Xiaomi’s benchmarks put it close to Opus 5 on several agent tasks. The deciding factor for local use will be hardware requirements and actual speed.
Introducing Xiaomi MiMo-V2.6 — Pro & Flash. Frontier intelligence, all the modalities, built in public. 🔹 Two omnimodal models, advancing through scaled reinforcement learning 🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks 🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models 🔹 Stronger coding, computer use, 3D reasoning and creative capabilities 🔹 Open model weights, technical report, RL environments and training code Blog:mimo.xiaomi.com/mimo-v2-6
7
1
30
2,272
I use 262144 tokens of total context per request and allow two active sequences, with a 1024 token prefill batch limit. Two is my current operating limit, not the hardware maximum. Extra requests queueI'm not claiming 4 full context agents stay fast.
2
653
The three default-off options I enable: GLM53_EXL3_MOE_FAST=1 GLM53_KDA_BF16_LARGE_M=1 GLM53_DENSE_FP8=dense,kda They target decode and prefill. PR233 adds ~3.3 GB/GPU. dense/KDA FP8 changes numerics. And remember, these are tradeoffs, not free speed.
3
1
9
826
Here's my current GLM 5.3 Flash EXL3 setup on 2x DGX Spark optimized for long coding sessions: with 3 opt-ins, 262k context and 2 active requests. My priority is responsive, stable long sessions, not the biggest context. Here's the profile and its tradeoffs 👇🧵
12
2
54
4,148
Replying to @rod_coutinho

ALT The Best Goat GIF by The Tonight Show Starring Jimmy Fallon

1
196
GLM 5.3 Flash EXL3 on 2x DGX Spark: ~15% faster long-prompt prefill! Our 100k token code test went from 71s to 62s. The update is live, but this optimization is experimental & opt-in. In a few hours, I’ll share which optimizations I actually enable and why. Update & more 👇
24
8
92
21,534