Cofounder & engineer. Running local AI, building things with it, and sharing what I learn.

New update merged for GLM-5.3-Flash EXL3 on 2× DGX Spark Loading the model now barely touches swap (99.99% less than before), long-prompt cache reuse is more reliable, and the stock 850k setup finally fits within its cache budget instead of overcommitting it, so it works out of the box. Update 👇
8
8
81
12,671
Heavy AI users will buy their own hardware once open models are good enough. I've burn 8B tokens a week consistently. At that volume, running models locally stops being a hobby and becomes a good complementary option. I'd bet that will happen for most in 2027. Hope there's stock left for everyone else by then...
72
dgpp, a C++/CUDA engine for DGX Spark, developed a system where each rank keeps its model slice resident on the GPU and boots from a per rank image cache in 15–30s, depending on the model. I just looked at adopting it, but our EXL3/DFlash2 and deepseek recipes aren't supported so far. Watching for compatible support...
4
28
1,528
Turns out I was hit by this hack. I lost my last valuable NFT, and not to the white hat, so it's probably gone for good. I don't see enough stories from people who actually lost their assets, so here's mine. Not every wallet got a rescue. Thanks a lot, Limit Break and Magic Eden…
Claim site is live. If I was able to save your NFTs, you can now reclaim them. You will have to revoke the PaymentProcessor approval first, if you haven't already. You may also opt to donate as part of your transaction. nftsaresafu.xyz/
3
6
913
On a DGX Spark, what you run still decides how fast it goes. For example, just turning on MTP in llama.cpp you can make Qwen3.8 27B on a single Spark far faster. Lots more cases like this everyday. We'll push this box to the limits until a new version is out...
6
2
30
1,753
Found this tab open on my Mac for the last 4 years. Back when GPT4 gave you 25 messages every 3 hours, and people went to Reddit asking how to get it back. @thsottiaux can we all get a banked reset for the loyalty I demonstrated, still subscribed?
2
5
566
Found myself asking Astra to tell me what it wants to say via Opus 5.5 lol That's how readable Opus is now.
Replying to @claudeai
Opus 5.5 communicates more naturally, addressing some of the most common feedback we heard on Opus 5. It puts the most important information up front and follows the writing rules you give it, which makes long sessions easier to follow.
1
8
1,163
NEW METHOD to register to Muse: Seen a lot of posts on opening a Muse account through cloud.browser-use.com. I tried it right away and it didn't work for me, looks like it got patched. A VPN didn't work either. What worked: open Google Gemini Spark, ask it to visit muse.ai/join and tell it to select “open via web”, or it redirects you to the homepage. Use an email that wasn't waitlisted. When the page opens, take control of the computer and enter your info manually. It won't accept typing the verification code for you. Then go back to the chat and tell it the birthday you want so it fills it in (I couldn't do that part from my phone). Finally connect your Facebook/Instagram account and you're done. If you get the chance, use my registration code PGPXC4 for 1B tokens to both of us.
1
5
695
Grok Bot is paying people for the templates they share on X! The pay depends on how many people use your @bot and how often, so useful ones earn the most. It could be a winning feature… pay for templates, the good ones show up, and the competitors are left behind.
2
3
582
Good Chinese openweight models will be optimized for Chinese hardware first. For DeepSeek V4s versions, Huawei Ascend are some of the only two stacks with optimized inference ready, alongside CUDA. The exceptions to this might be the small models that fit gaming GPUs.
2
1
20
2,136
New update merged for GLM-5.3-Flash EXL3 on 2× DGX Spark Loading the model now barely touches swap (99.99% less than before), long-prompt cache reuse is more reliable, and the stock 850k setup finally fits within its cache budget instead of overcommitting it, so it works out of the box. Update 👇
8
8
81
12,671
For this to take effect for now you need to use the auto loader (LOAD_FORMAT=), model loading will barely touch swap. In the coming days aside from improving decode, prose and prefill we'll look into standardizing and simplifying all the options available. github.com/MiaAI-Lab/GLM-5.3…
1
13
857
OpenCode's data pages already list Kimi K4, GLM 5.5 Flash, DeepSeek V4.1 Pro and Qwen 3.8 Max Preview! They're placeholder entries, all four show 0% usage and most of their specs say unknown. Pacing the frontier isnt working, and Kimi K4 is the one I want to try first!
🚨惊了!OpenCode 数据页提前曝光下一批模型,目录已经挂上! 这不是官宣能用,是模型 ID 已经进库: 🔹Kimi K4(Moonshot)
opencode.ai/data/moonshot/ki… 🔹GLM 5.5 Flash(Zhipu)
opencode.ai/data/zhipu/glm-5… 🔹DeepSeek V4.1 Pro
opencode.ai/data/deepseek/de… 🔹Qwen 3.8 Max Preview Free
opencode.ai/data/qwen/qwen3-… 🔹Muse Spark 1.4 Contributor(Meta)
opencode.ai/data/meta/muse-s… 🔹额外同批出现:
Qwen 3.8 Max Prime(列表标注 9/23,暂无用量)
opencode.ai/data/alibaba 现役还是 K3 / GLM-5.3-Flash / V4.1 Flash。K4、5.5、V4.1 Pro 才是下一代信号。 #OpenCode #AI #KimiK4 #GLM55 #DeepSeek #Qwen #MuseSpark #OpenSource
4
22
2,642
Two big crypto hacks back to back: - Bitget says the attackers faked transaction data and got its own approval system to sign off on about $352M in withdrawals. - While an NFT payment contract let someone take thousands of NFTs people had approved to it, for free. I'd bet human run AI teams are behind hits like this. Nobody can skimp on security anymore, and still so many do...
2
8
858
Got access to the GLM 5.3 FlashX. It's the same weights as 5.3 Flash, just served at 100 to 200 tok/s, for 2.5x the quota. So it's not smarter and you're paying 2.5x for speed alone... dont you think 100-200 tok/s should be the standard for flash on the cloud? why a premium!?
2
24
1,817
Got temporary access to a Mac Studio M5 Ultra. Barely optimized, Qwen 3.8 Flash Next runs at 85.9 tok/s and 6085 tok/s prefill on long prompts. A 2x DGX Spark recipe gets 52.1 tok/s single stream and about 2960 tok/s prefill on 16k–64k prompts. Already ahead of 2x Spark, but I was expecting a much much better baseline…
26
5
122
40,075
A $500 month @OpenAI plan showed up in ChatGPT's code, likely running on @cerebras with higher limits. That could mean up to 750 tok/s, the speed Cerebras claimed for 5.6 Sol. You'd use up your limits as fast as Pro, just with way more tokens used. Who's getting this?
1
5
1,768
Left for a business trip 2 days ago. Bought an Aqara smart plug for my whole computing cluster, it arrived the night before I left… and it's too big for my wall socket. So now I'm just hoping nothing OOMs while I'm away, with a lot of EXL3 upgrades still in testing…
7
11
1,082
ZAI got caught with ZCode uploading users repos, open sourced it, and now they are giving a free quota reset hidden in the app (30 days to use it) + 300M free GLM 5.3 Flash tokens this weekend. Only claimable inside the app that got caught, but I'll take it…
3
14
1,082
Think of the possibilities: a self driving car practicing a packed intersection where every driver and pedestrian is an agent. Or a robot learning to work around people who never get tired of testing it. It could have bigger potential than multiplayer… why would they focus on that?
Introducing Agora-2, our next-generation multi-agent world model. Agora-2 supports up to 20 humans and agents interacting inside a shared environment, all simulated in real time. Our multiplayer research preview is available to try right now!
4
1,226