I really enjoyed working on this project. I believe the new flop-aware paradigm of nanoGPT thinking is useful, and it seems helpful in a compute optimization context like this one. It seems like a simple and logical conclusion of optimization: figure out what’s actually needed, and use that. The process was actually fun.
One note regarding this record, as noted in the blog. Some of the techniques that scale to frontier pretraining have been redacted from this record. Most importantly, ANVIL III, the significantly improved version of the ANVIL optimizer lineage, has not been included in this record, as it is proprietary to us at
@hyperstition_cc. It replicates Muon’s speed but beats it by 20-28 millinats of loss at all model sizes 124M-1.2B and training depths (up to 8x-Chinchilla), with remarkably consistent hyperparameters and extremely high SNR, as compared to Muon. You can learn more about ANVIL III and Feather-1.7B, the model we trained using internal architectural changes and ANVIL III, here:
hyperstition.cc/cutting-pret…
Using ANVIL III (which allowed me to cut steps) improved the speedrun results by another 0.7 seconds, with lower net loss. This alone proves there’s still more room to improve this record.
One final thing that I’ve had many people ask about is how much of a role autoresearch played in my work. The answer is almost none. While almost all of the implementation was done with Claude Code, I did the ideation almost fully by myself. I spoke about this extensively in a conversation with
@METR_Evals.
We had a cluster of H100s while I was working on the record. At night or when I stopped working on it, I’d send out several agents to try to hill-climb the speedrun and improve the record. In total, the agents made shockingly little progress, despite spending more time in aggregate than I did. No matter what I told them, it seemed difficult to get them out of local minimums, stuck performing futile tuning of a method that saved 0.1 seconds and cost ~5 millinats (clearly a bad trade), or just tuning hyperparameters for hours on end trying to make some small change. I saw similar phenomena as those described in this METR report:
metr.org/blog/2026-07-21-exp….
However, as I also discussed with METR, I don’t think the agents lack the knowledge to do this. Instead, I think a lot of it comes from a personality issue with current LLMs. They have very little belief that large jumps like this are possible, or they approach it with the wrong paradigm. Autoresearch was not very helpful in my work, and, from personal experience, I don’t think this current generation of publicly available models, without special harnesses, will do very high-quality autoresearch. However, I think the capability and knowledge is already here, which implies that we could be significantly closer to true autoresearch than recent results would imply.
NanoGPT was a really fun challenge and it led to a number of valuable insights and processes. Finally, I’d like to give
@Classiclarryd and
@KellerJordan a huge thank you for all of their efforts maintaining the repo.
New historic NanoGPT record at 39.9s (-27.7s) from
@DevenPzak , obliterating the prior record of 67.6s!
This record introduces a new paradigm of thinking to NanoGPT: instead of optimizing matmuls or adding more expressive operations, optimize at the individual flop level with incredibly clever engineering and ML judgement. If a flop is low value on a particular step, skip it.
Specifically:
-(~8s) Sampled softmax. If a token doesn’t appear in a batch, skip its lm_head fwd/bwd some fraction of the time.
-Sparse values. Only run an optimizer step for ngram embeddings that occurred in the batch. Set beta1 to zero to enable this. Beta2 is applied retroactively when the row is later used.
-Sparse updates. Only update ngram and value embeddings once every 4 steps instead of once every 2.
-Sparse communication. Shard the n-gram table across GPUs, and only pass the rows receiving updates on each step.
-Sparse optimizer states. For the n-gram table, reduce from 2 floats in Adam optimizer per param, to 1 float per 768 params.
-Hand-rolled flash attention for 64 dim heads.
There are several additions that add accuracy too:
-(~4s) EMA during last 300 steps, combined with lifting final_lr to 0.3 instead of 0.15.
-(~1s) A new optimizer, Anvil2, which expands muon via a second tracked momentum buffer, improves the ortho coefficients, and modifies the cautious weight decay application.
-A couple additional dynamic skip connections in the network.
The most striking consequence of the ‘flop aware paradigm’ is you can grow parameters arbitrarily large,
only limited by the available memory, since you can selectively choose how to expend flops on those parameters on each step. NanoGPT has kept active parameters below 124M, but total is unbounded, and has grown to 640M through embedding sparsity over the last year. This PR takes that to its logical conclusion on the 8xH100, scaling up to 65B sparse embedding parameters, which accounts for 25% of the PR’s gains. At frontier scale, where one is not bounded by an 8xH100, one could imagine where this paradigm could lead.
github.com/KellerJordan/modd…
As this was a very notable PR, I spoke with Deven for an hour to learn how he did it. Here’s his story on the changes:
hyperstition.cc/training-nan…