I build the AI systems I use to run GTM. Previously 0x, AWS, OpenSea. A running log: what I built, what broke, what surprised me. No hype.

RAG doesn’t require the model to memorize an enterprise’s knowledge; it retrieves the right knowledge at inference time, while embeddings and rerankers determine whether the model gets the right context in the first place. Today we move up the stack and answer a question that appears constantly in enterprise AI. Link to post in comments 👇
1
48
7
70
William Doom retweeted
Industrial software doesn’t have to feel like industrial software Manage warehouses like you’re playing a strategy game Built with Opus 5.5
640
2,082
24,421
2,224,007
William Doom retweeted
Lots of inventive things made with Claude these past few weeks. A few of our favorites: A working watermill, built with Opus 5.5.
The Bugged Dev
359
666
10,823
755,029
AEO
Vercel is generating $600 million in annualized revenue, up 148% from a year earlier, as coding agents become a major source of new business. Agents now account for about half of new business, up from less than 3% at the start of the year. Read more details: thein.fo/4eaW4bv
3
71
William Doom retweeted
Born in Higgsfield. All over your feed. Introducing Higgsfield AI Influencer. Create your own AI influencer and bring them into any trend. Try now with up to 5 FREE generations on Higgsfield and in the ChatGPT extension. Powered by Genjutsu. Credits: @JeanPhilMadame
589
520
6,852
3,307,875
William Doom retweeted
As a JAY-Z fan all my life; 1. I caught the “keep 1 eye open like CBS” line immediately. I did NOT register that he also said “SEE BS”. 2. I didn’t know the song was about paranoia, from the 1st bar to the hook, to the last bar of the song. All about paranoia.
93
499
2,728
278,655
Funny, but this is the real lesson, in AI infra, the rep who can answer the KV cache question is the one who wins the deal. I've been writing up the technical side of selling inference so GTM folks don't get caught at the booth itsthedoom.substack.com/p/ba…
1
5
94
How about, just don’t give your kids a phone
A phone that gets kids OUTSIDE and looking up instead of down. Yes, please. I pre-ordered mine (along with thousands of people in the last 24hrs).
1
1
83
I’ve started thinking about quantization less as model compression and more as a capacity lever. If you can reduce the memory footprint while preserving acceptable quality, that potentially changes how much expensive GPU infrastructure the workload requires. link to post in comments 👇
1
49
Back 2 back
33
William Doom retweeted
Year 1 vs Year 24
109
2,510
33,227
1,150,222
49
One observations about inference economics is that some of the biggest gains aren’t necessarily about cheaper GPUs. They’re about eliminating unnecessary computation and extracting more useful work from the compute you’re already paying for. Batching gets more useful work from each GPU; caching avoids making the GPU repeat work it has already done. Both turn serving architecture directly into inference economics. What if you could make a model require substantially less expensive compute, but doing so might slightly change its quality? full post in comments 👇
Made with AI
1
41
William Doom retweeted
Get ready.
2,984
3,728
64,322
17,413,763