billions will thrive robot inference @hf0, ex-quant 2x

SF | Montreal
Pinned Tweet
I’m hiring for my robotics inference startup, we’re working on solving fast cheap long context batch invariant Inference for VLAs and WMs, billions will thrive!
43
20
401
23,169
Lucas retweeted
I really enjoyed working on this project. I believe the new flop-aware paradigm of nanoGPT thinking is useful, and it seems helpful in a compute optimization context like this one. It seems like a simple and logical conclusion of optimization: figure out what’s actually needed, and use that. The process was actually fun. One note regarding this record, as noted in the blog. Some of the techniques that scale to frontier pretraining have been redacted from this record. Most importantly, ANVIL III, the significantly improved version of the ANVIL optimizer lineage, has not been included in this record, as it is proprietary to us at @hyperstition_cc. It replicates Muon’s speed but beats it by 20-28 millinats of loss at all model sizes 124M-1.2B and training depths (up to 8x-Chinchilla), with remarkably consistent hyperparameters and extremely high SNR, as compared to Muon. You can learn more about ANVIL III and Feather-1.7B, the model we trained using internal architectural changes and ANVIL III, here: hyperstition.cc/cutting-pret… Using ANVIL III (which allowed me to cut steps) improved the speedrun results by another 0.7 seconds, with lower net loss. This alone proves there’s still more room to improve this record. One final thing that I’ve had many people ask about is how much of a role autoresearch played in my work. The answer is almost none. While almost all of the implementation was done with Claude Code, I did the ideation almost fully by myself. I spoke about this extensively in a conversation with @METR_Evals. We had a cluster of H100s while I was working on the record. At night or when I stopped working on it, I’d send out several agents to try to hill-climb the speedrun and improve the record. In total, the agents made shockingly little progress, despite spending more time in aggregate than I did. No matter what I told them, it seemed difficult to get them out of local minimums, stuck performing futile tuning of a method that saved 0.1 seconds and cost ~5 millinats (clearly a bad trade), or just tuning hyperparameters for hours on end trying to make some small change. I saw similar phenomena as those described in this METR report: metr.org/blog/2026-07-21-exp…. However, as I also discussed with METR, I don’t think the agents lack the knowledge to do this. Instead, I think a lot of it comes from a personality issue with current LLMs. They have very little belief that large jumps like this are possible, or they approach it with the wrong paradigm. Autoresearch was not very helpful in my work, and, from personal experience, I don’t think this current generation of publicly available models, without special harnesses, will do very high-quality autoresearch. However, I think the capability and knowledge is already here, which implies that we could be significantly closer to true autoresearch than recent results would imply. NanoGPT was a really fun challenge and it led to a number of valuable insights and processes. Finally, I’d like to give @Classiclarryd and @KellerJordan a huge thank you for all of their efforts maintaining the repo.
New historic NanoGPT record at 39.9s (-27.7s) from @DevenPzak , obliterating the prior record of 67.6s! This record introduces a new paradigm of thinking to NanoGPT: instead of optimizing matmuls or adding more expressive operations, optimize at the individual flop level with incredibly clever engineering and ML judgement. If a flop is low value on a particular step, skip it. Specifically: -(~8s) Sampled softmax. If a token doesn’t appear in a batch, skip its lm_head fwd/bwd some fraction of the time. -Sparse values. Only run an optimizer step for ngram embeddings that occurred in the batch. Set beta1 to zero to enable this. Beta2 is applied retroactively when the row is later used. -Sparse updates. Only update ngram and value embeddings once every 4 steps instead of once every 2. -Sparse communication. Shard the n-gram table across GPUs, and only pass the rows receiving updates on each step. -Sparse optimizer states. For the n-gram table, reduce from 2 floats in Adam optimizer per param, to 1 float per 768 params. -Hand-rolled flash attention for 64 dim heads. There are several additions that add accuracy too: -(~4s) EMA during last 300 steps, combined with lifting final_lr to 0.3 instead of 0.15. -(~1s) A new optimizer, Anvil2, which expands muon via a second tracked momentum buffer, improves the ortho coefficients, and modifies the cautious weight decay application. -A couple additional dynamic skip connections in the network. The most striking consequence of the ‘flop aware paradigm’ is you can grow parameters arbitrarily large, only limited by the available memory, since you can selectively choose how to expend flops on those parameters on each step. NanoGPT has kept active parameters below 124M, but total is unbounded, and has grown to 640M through embedding sparsity over the last year. This PR takes that to its logical conclusion on the 8xH100, scaling up to 65B sparse embedding parameters, which accounts for 25% of the PR’s gains. At frontier scale, where one is not bounded by an 8xH100, one could imagine where this paradigm could lead. github.com/KellerJordan/modd… As this was a very notable PR, I spoke with Deven for an hour to learn how he did it. Here’s his story on the changes: hyperstition.cc/training-nan…
6
17
146
7,011
All of inference engineering is heading in this direction, create envs per model and chip spec and loop to tpsmaxx If you are interested building something like this, our harness is for making VLA/WM inference much faster!
2.94× KDA fwd, 6.84× KDA bwd, 2.59× MSA prefill, 3.99× MSA decode, 1.33× KDA decode, 1.71× MLA, 1.68× VSA — just by pressing Enter. After years of working on ML compilers and designing many abstractions, the most useful compiler thing in TIRx I’ve ever built might just be a PTX tablegen.
1
4
6,535
Lucas retweeted
i think modded-nanogpt is literally the future of research
1
16
490
is anyone working on interconnects between b200, mi355x & TPU/ xpu
2
1
4
304
Lucas retweeted
We are releasing Dyna-2.1, the first Physical Agent that achieves reliable super long-horizon whole-body autonomy. It combines our brand-new semi-humanoid hardware with an agentic system built around Dyna-2 to handle ultra-long real-world workflows. Here is an uncut footage of Dyna-2.1 completing an entire hour-long laundry room workflow, just like a human does.
34
145
787
174,260
is there any interest in having humans operators behind our api so that we can bring robots to prod faster instead of waiting for autonomy
3
7
1,019
Lucas retweeted
We instituted an official AI Writing Policy at @Clay. Massive thanks to @sophiebits on our engineering team who wrote this. Originally this was just for eng, but other teams found it so helpful we expanded it company-wide. Here are the four guiding principles: 1. You must stand behind every idea and sentence It is your responsibility to make sure that the entire document is representative of your own thoughts before you share it. 2. Writing is thinking Spending time on the writing process teaches you more about your topic. If you circumvent this process, you will walk away with a poorer understanding of the subject matter. 3. More time should be spent writing a document than consuming it If you generate a document from a short prompt then ask your readers to go through the longer output, you are disrespecting their time. They can talk to ChatGPT themselves if they want to. 4. Longer is not better AI makes it easy to generate long docs, and it loves padding them with sentences that say nothing. If you're producing docs from a short prompt, consider just sharing the prompt. You can read (and borrow) the full policy below. Hopefully we all communicate a bit more clearly now :)
169
345
3,963
675,653
At batch 10, a b300 can sever dreamZero Wam at 30hz, with 75% cache hits, it can serve ~110–145 reactive robots. current spot pricing is $8/h. useful robot work is currently at $0.10/h, the world is not ready for this abundance of labor
1
25
7,934
Life is for sowing, the harvest isn’t here yet
1
4
359
Lucas retweeted
when you sell infra to startups there is a venture investing quality to it if your clients can't win & grow fast over the next years, you're ngmi but if you're useful early to people about to win, you're having the easiest sell process for the best future outcome at boat.dev we optimize everything from product positioning, gtm and tech for that, maximally serving the future winners - be where they are: on X - be how they are: opinionated, outspoken, grinding, high-integrity - optimize for what they'll do but large incumbentd don't: infra for long-running employee agents - sell to who decides for them: their agents (publish lots of data, benchmarks, have detailed docs)
6
2
46
2,282
yo big dawg, i could getchu 12 B300 nodes no crap fr. $4.50/hr, 5 year term, 30% upfront, deliver in 12 weeks ima need you to wire me $5.6 milli by monday or gtfo no i've never done this before. here's some info from claude and naval's spf is an investor so you know im legit
28
10
316
25,583
Lucas retweeted
the best people to work with are: - high agency - high ownership - low ego / hungry for feedback
45
108
1,388
33,299
Lucas retweeted
The unreasonable effectiveness of actually looking at data
We (@PantheonInc) have been working on a robotics data quality pipeline that uncovered a series of major problems in public robotics datasets, especially for world modeling. To improve the quality of data available to open-source robotics, we're publishing annotations for four of the most popular datasets. Some examples of issues, and our report 🧵
2
3
36
26,620
I spoke to some MBAs recently, and this was my provocation: You are here to become big important CEOs some day... but given modern AI tools, you kinda already are one! Everyone has a company of smart people in their pocket now. So why wait? Start acting like a CEO now!
It's a little hard to internalize that the list of things one person can achieve is rapidly growing. You must think big to match what you're truly capable of now
13
8
219
14,691
Lucas retweeted
Just got revenge on someone who wronged me 6 years ago. Never be relaxed ever. I'm coming
63
1,731
6,704
136,082
I’m going to make a 20$ physical Ai inference sub with the strict purpose to just try to break even. At scale I think it’s possible
1
1
17
804
400 years of human evolution in 1 photo
770
15,206
220,970
12,053,588
deployments data hardware compute financing
1
26
1,442
“I have a great cluster available for you perfect for inference” > shit interconnects
2
13
1,022