Building @zml_ai (and we're hiring), ex @zenly, ex Exalead, ex @google. Skydiver and wingsuiter.

Paris, France
6
2
48
Great read on developing with AI
I've got an agent in a loop optimizing a renderer with the goal to minimize frame times (and tests to measure). It got times down from 88ms to 2ms and allocations down from ~150K to 500. Sounds good, right? Wrong. This is exactly why agent psychosis is a big fucking problem. As an experiment, I rewrote the Ghostty core render state in Go, with access to identically laid out data structures as Ghostty and the exact same validation tests. I made a purposely naive renderer (simple, correct, but slow). 88ms per frame with 150,000 allocations (horrendous, lol)! I then kickstarted a Ralph loop to bring the frame times down. I told it it can't modify input data structures or the public API or tests (they're correct), but it can do anything else it wants. It got to work. It has worked for about 4 hours. I've spent around $350 on this experiment so far. The results? 88ms => 1.5ms 150K allocs => ~500 allocs Incredible right? Nope. My hand-written renderer I ported has frame times (same benchmark) of ~20us (0.020ms) and 0 allocations in the update path. This is the problem with psychosis and lacking systems understanding. If you don't understand the system, you're going to accept that this is an incredible result. If you understand the system, you'll see better solutions immediately and can do roughly 75x better on throughput. The people who blindly trust agent output are in the former camp. They're sheeple, overdrinking from a fountain of mediocrity. Standard disclaimer: I use AI all the time. I like AI. The point I'm making is to not blindly accept results. Think. Analyze. Learn.
12
1,693
Steeve Morin retweeted
Oh come on. Quick look at colorshift (13x faster): HashMap in hot loop (unoptimized Rust vibeslop) is slower than indexed arrays (ASM). No shit. Numbers like 17x should raise eyebrows but reading the comments nobody is thinking anymore. This AI psychosis has fried people's brains
I ported the Omarchy screensaver engine (ttfx) from Rust to x86-64 assembler, and it's up to 17x faster!! One-shot translation by Opus 5.5. We keep drilling until the agentic drill bit hits bedrock! github.com/omacom/ttfx/pull/…
57
114
2,565
237,401
it's live (if you know where to look) but before i blog there's a few numbers i need to check
big release tomorrow
1
11
2,208
also, where can we post our benchmarks ?
5
535
i'm super serious, who has the fastest Deepseek 4.1 Flash with the actual batch size ? i want to double check our numbers
4
9
2,125
👏👏👏👏👏👏👏👏
HOTEL LOBBY but the beat is Hotel California
2
840
i was told $8 would fix it though
The 300 “people” in your replies calling you a traitor because you don’t want to drown refugees
4
817
the 45 minutes I've spent trying to run DS 4.1 Flash on this GPU make me conclude one thing: day 0 my ass
1
6
1,584
few understand how IOCP mogged
Also it got I/O Completion Ports (IOCP) quite a bit before Linux got io_uring I believe. I used to get roasted by Linux folks when I said the I/O kinda sucked compared to the NT kernel. You can be a fan of the NT kernel even if you don’t like the direction of Windows.
2
633
meta ditching the llama brand for "muse" (whatever this is) is a generational blunder
this is like steve jobs unveiling the iphone
4
20
5,780
yeah so I was pissed there were NO tpus to rent, but now i'm *super* pissed
Can our TPUs survive and operate in space? Well, we're going to find out. Project Suncatcher is hitching a ride aboard @SpaceX's Transporter-18 mission, testing a prototype satellite built in partnership with @planet One small step for TPUs....
5
17
2,553
big release tomorrow
2
1
64
4,479
if you have been sleeping, you shouldn't sadly, there's nothing you can do because there's literally no TPUv7 to get according to Google ("fully booked for a few weeks at least")
Btw, TPU support on zml/llmd is looking very, very good. We have it working, paged attention and all. It’s a *lot* faster than alternatives. I’ll try to make a video but I didn’t think 4x v5e could spit out that many tok/s (in bs=1 at least). Hope to release it soon!
1
1
18
2,376
gentle reminder btw
TPU V7s are fully booked for the "next few weeks", like, it's impossible to get any
1
9
2,091
TIL DEBUG_HIP_FORCE_GRAPH_QUEUES=1 🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️🤦‍♂️
2
609
ouch
Replying to @dylan522p
Your update just killed your original accusation. You said AMD had found a business opportunity “dumping US military chips” into China at 1/4 the U.S. price. Now AMD tells you it did not sell or ship this XCZU47DR to Puzhi. If that is true, AMD was not the seller setting the alleged $1K China price. Your dumping theory collapses into a provenance investigation. And there is an obvious paper trail to pull. Puzhi’s exact P047 is publicly listed as part of AMD’s own FPGA Playground, while the Playground advertises discounted AMD parts pricing for qualifying initial production runs. The finished P047 is currently sold by Crowd Supply for $8,749 worldwide. (Crowd Supply) So ask AMD for the useful things: lot/date codes → authorized distributor → invoice → export/re-export authorization → consignee → end user → program subsidy/discount records → whether the devices are genuine, diverted, secondary-market, reclaimed, or counterfeit. The part itself is publicly classified by Avnet as ECCN 3A001.A.14, so provenance and licensing are real questions. (Avnet) AMD itself says export controls govern FPGA transfers and downstream re-exports. (AMD) “AMD says they’re looking into it” is not vindication. It is the point where an analyst starts working. And “I’ve heard there are bitstream-compatible non-AMD FPGAs floating around” is a completely separate counterfeit/clone/provenance hypothesis. Naming another hypothesis after the first one broke does not repair the first one. You began at AMD treason/dumping, got corrected into unknown sourcing, and are now presenting AMD opening an internal ticket as confirmation that Perfect Dylan was right all along. Pull the receipts. Then write the story.
1
5
2,322
our obsession with zero python is getting us in great places
9
5
77
7,342
now Qwen 3.8 27B BF16 19500 tok/s at bs=1024 on 4xGB300
18300 tok/s at bs=1024
3
34
1,802