Trying to solve sample efficiency prev: Engineering physics @iitbombay '23

44% on ARC-AGI-1 in 67 cents! Trained from scratch in 2hrs on a 5090 Matches TRM, beats HRM and is way faster & cheaper No recursion, just a transformer Also, 7% on ARC-2 🧵
31
76
768
64,260
Mithil Vakde retweeted
I think I just trained NanoGPT in 9.65 seconds!? Yesterday, @hermanbrunborg used a connected longest exact match model, trained concurrently on CPU, to get the nanoGPT record from 39.9 seconds to 21.5 seconds, in an impressive open PR. By adapting their work, and generalizing the CPU model to a larger share of the tokens, I got the record down to 9.65 seconds, and submitted the PR this morning. I built on three open PRs, in addition to my contributions: the main basis for my generalization was from @hermanbrunborg, and additional work from @nthngdy, @cyrusghane, and Daniel Monroe. Their engineering is extremely impressive. Of course, all of our records are still subject to review. As I mentioned in my comment on Herman‘s PR (and in my own PR description), CPU usage is a gray area in nanoGPT. Their PR was the first to discover a clever way to effectively use CPU during the run, so I’m looking forward to the official reviews. This was another very fun challenge! Particularly (potentially) breaking the 10 second barrier yesterday. So much engineering tact is yet to be explored. Huge shoutout to the three authors of the PRs for their innovations. @hermanbrunborg - I will also be at stanford for the weekend along with some of my cofounders, and sf for a couple days after, would love to chat more IRL, if you’d be open 🙂
20
40
634
55,103
engineering physics gives you the best of both worlds :)
Replying to @DavidDeutschOxf
The future light cone of discoveries rounds up to 100% engineering, as the set of physics rules that govern our reality is very tiny compared to what can be done with those rules. The reason I didn’t choose a career in physics, my favorite subject, is that I would be stuck waiting for a new collider or telescope. Better to advance technologies that ultimately enable discovery of new physics.
10
1,417
Mithil Vakde retweeted
After seven years, I've made the hard decision to leave Google Deepmind. I joined Google in 2019 because I believed we could make build models to understand and generate code. Back then, very few people had that conviction. I’m glad I found those people at Google and that we worked hard to make that dream a reality. Coding is now democratised and it will be used to solve everything else!
80
15
578
46,665
Mithil Vakde retweeted
Can complex multi-step reasoning emerge purely from cells that only talk to immediate neighbors? Happy to share our paper “Reasoning with Neural Cellular Automata (NCA)”, from our team at Google, Paradigms of Intelligence 🧵👇
45
181
1,255
129,825
Mithil Vakde retweeted
My second interview with @CJHandmer, Founder of @TerraformIndies. 0:09 How the Mass Driver is going to work 2:54 Mass Driver vs scaling Starship 4:26 Economics of space compute vs building data centers on Earth 8:53 Why large scale projects are becoming illegal 14:26 Terafab 20:02 Starmind 26:19 Maximum scale as fast as possible 27:43 Luisiana launch tower buildout 30:13 Elon’s decade long company baking period 32:56 Firing teams 37:35 Starting companies 41:57 Being unfiltered publicly 42:52 DOGE 46:39 Deregulating without a major war 55:17 Founder-led companies are mini dictatorships 58:53 Taxes 1:03:26 Civilization hard resets 1:05:30 Starship scaling constraints 1:07:16 How Elon sets objectives 1:09:31 Universal high income
23
50
369
64,674
the only way to ensure long term survival of human species is building superintelligence if no one builds it, our species will die
imo one of the best arguments in favor of building superintelligence is so we can point it at FTL drives. maybe worth going slow since we probably have at least hundreds of years, but in the long term I see FTL x-risk as comparable to AI x-risk.
11
995
Mithil Vakde retweeted
I really enjoyed working on this project. I believe the new flop-aware paradigm of nanoGPT thinking is useful, and it seems helpful in a compute optimization context like this one. It seems like a simple and logical conclusion of optimization: figure out what’s actually needed, and use that. The process was actually fun. One note regarding this record, as noted in the blog. Some of the techniques that scale to frontier pretraining have been redacted from this record. Most importantly, ANVIL III, the significantly improved version of the ANVIL optimizer lineage, has not been included in this record, as it is proprietary to us at @hyperstition_cc. It replicates Muon’s speed but beats it by 20-28 millinats of loss at all model sizes 124M-1.2B and training depths (up to 8x-Chinchilla), with remarkably consistent hyperparameters and extremely high SNR, as compared to Muon. You can learn more about ANVIL III and Feather-1.7B, the model we trained using internal architectural changes and ANVIL III, here: hyperstition.cc/cutting-pret… Using ANVIL III (which allowed me to cut steps) improved the speedrun results by another 0.7 seconds, with lower net loss. This alone proves there’s still more room to improve this record. One final thing that I’ve had many people ask about is how much of a role autoresearch played in my work. The answer is almost none. While almost all of the implementation was done with Claude Code, I did the ideation almost fully by myself. I spoke about this extensively in a conversation with @METR_Evals. We had a cluster of H100s while I was working on the record. At night or when I stopped working on it, I’d send out several agents to try to hill-climb the speedrun and improve the record. In total, the agents made shockingly little progress, despite spending more time in aggregate than I did. No matter what I told them, it seemed difficult to get them out of local minimums, stuck performing futile tuning of a method that saved 0.1 seconds and cost ~5 millinats (clearly a bad trade), or just tuning hyperparameters for hours on end trying to make some small change. I saw similar phenomena as those described in this METR report: metr.org/blog/2026-07-21-exp…. However, as I also discussed with METR, I don’t think the agents lack the knowledge to do this. Instead, I think a lot of it comes from a personality issue with current LLMs. They have very little belief that large jumps like this are possible, or they approach it with the wrong paradigm. Autoresearch was not very helpful in my work, and, from personal experience, I don’t think this current generation of publicly available models, without special harnesses, will do very high-quality autoresearch. However, I think the capability and knowledge is already here, which implies that we could be significantly closer to true autoresearch than recent results would imply. NanoGPT was a really fun challenge and it led to a number of valuable insights and processes. Finally, I’d like to give @Classiclarryd and @KellerJordan a huge thank you for all of their efforts maintaining the repo.
New historic NanoGPT record at 39.9s (-27.7s) from @DevenPzak , obliterating the prior record of 67.6s! This record introduces a new paradigm of thinking to NanoGPT: instead of optimizing matmuls or adding more expressive operations, optimize at the individual flop level with incredibly clever engineering and ML judgement. If a flop is low value on a particular step, skip it. Specifically: -(~8s) Sampled softmax. If a token doesn’t appear in a batch, skip its lm_head fwd/bwd some fraction of the time. -Sparse values. Only run an optimizer step for ngram embeddings that occurred in the batch. Set beta1 to zero to enable this. Beta2 is applied retroactively when the row is later used. -Sparse updates. Only update ngram and value embeddings once every 4 steps instead of once every 2. -Sparse communication. Shard the n-gram table across GPUs, and only pass the rows receiving updates on each step. -Sparse optimizer states. For the n-gram table, reduce from 2 floats in Adam optimizer per param, to 1 float per 768 params. -Hand-rolled flash attention for 64 dim heads. There are several additions that add accuracy too: -(~4s) EMA during last 300 steps, combined with lifting final_lr to 0.3 instead of 0.15. -(~1s) A new optimizer, Anvil2, which expands muon via a second tracked momentum buffer, improves the ortho coefficients, and modifies the cautious weight decay application. -A couple additional dynamic skip connections in the network. The most striking consequence of the ‘flop aware paradigm’ is you can grow parameters arbitrarily large, only limited by the available memory, since you can selectively choose how to expend flops on those parameters on each step. NanoGPT has kept active parameters below 124M, but total is unbounded, and has grown to 640M through embedding sparsity over the last year. This PR takes that to its logical conclusion on the 8xH100, scaling up to 65B sparse embedding parameters, which accounts for 25% of the PR’s gains. At frontier scale, where one is not bounded by an 8xH100, one could imagine where this paradigm could lead. github.com/KellerJordan/modd… As this was a very notable PR, I spoke with Deven for an hour to learn how he did it. Here’s his story on the changes: hyperstition.cc/training-nan…
10
31
267
19,209
man just rip the bandaid off, no more UIs make it agents all the way down auth and info/entertainment are the only time I wanna see an interface
5
22
1,720
Mithil Vakde retweeted
New historic NanoGPT record at 39.9s (-27.7s) from @DevenPzak , obliterating the prior record of 67.6s! This record introduces a new paradigm of thinking to NanoGPT: instead of optimizing matmuls or adding more expressive operations, optimize at the individual flop level with incredibly clever engineering and ML judgement. If a flop is low value on a particular step, skip it. Specifically: -(~8s) Sampled softmax. If a token doesn’t appear in a batch, skip its lm_head fwd/bwd some fraction of the time. -Sparse values. Only run an optimizer step for ngram embeddings that occurred in the batch. Set beta1 to zero to enable this. Beta2 is applied retroactively when the row is later used. -Sparse updates. Only update ngram and value embeddings once every 4 steps instead of once every 2. -Sparse communication. Shard the n-gram table across GPUs, and only pass the rows receiving updates on each step. -Sparse optimizer states. For the n-gram table, reduce from 2 floats in Adam optimizer per param, to 1 float per 768 params. -Hand-rolled flash attention for 64 dim heads. There are several additions that add accuracy too: -(~4s) EMA during last 300 steps, combined with lifting final_lr to 0.3 instead of 0.15. -(~1s) A new optimizer, Anvil2, which expands muon via a second tracked momentum buffer, improves the ortho coefficients, and modifies the cautious weight decay application. -A couple additional dynamic skip connections in the network. The most striking consequence of the ‘flop aware paradigm’ is you can grow parameters arbitrarily large, only limited by the available memory, since you can selectively choose how to expend flops on those parameters on each step. NanoGPT has kept active parameters below 124M, but total is unbounded, and has grown to 640M through embedding sparsity over the last year. This PR takes that to its logical conclusion on the 8xH100, scaling up to 65B sparse embedding parameters, which accounts for 25% of the PR’s gains. At frontier scale, where one is not bounded by an 8xH100, one could imagine where this paradigm could lead. github.com/KellerJordan/modd… As this was a very notable PR, I spoke with Deven for an hour to learn how he did it. Here’s his story on the changes: hyperstition.cc/training-nan…
37
156
1,468
414,858
Mithil Vakde retweeted
you think thats a faithfully verbalized chain of computation you're reading?
3
42
550
11,901
Mithil Vakde retweeted
After 2.5 years at Sarvam, I am starting something new! I had a great time leading the RL team. We launched the 30B and 105B models back in Feb with a young team building everything from scratch. Excited to see the team keep growing and pushing the frontier. Grateful to @pratykumar for the early bet and trust. As for what's next: I am convinced that there are a few critical parts of the stack that we still need to build to meaningfully accelerate science. One of them is verification. Scientific discovery only scales with compute when proposing and verifying new ideas is fast and cheap. In math and code, we're starting to see what happens when models can generate ideas and reliably check them at scale. OpenAI's recent Navier-Stokes work is a glimpse of what that looks like. Similar moments in chemistry, materials, biology and the rest of the physical sciences would require verification to get dramatically faster and cheaper. It cannot be bottlenecked by the throughput of physical labs More soon. Very excited for what's next.
101
64
1,994
233,830
Mithil Vakde retweeted
today we’re launching eyecandy robotics! we’re bringing animated characters to life through robotics and ai.
143
119
1,373
166,083
Models not having "taste" is a very short sighted criticism imo it all depends on where the optimisation pressure is. As we move rewards up the abstraction ladder taste will develop Eg: reward the model during postraining for finding better NN architectures for some X goal. Initially it will do dumb brute force search. But eventually it will learn shortcuts which we'll call taste
3
30
1,213
dw babe the singularity is gonna be fun
I made this with one prompt using Opus 5.5 I spoke to my computer for 5mins, claude worked for 12 hours, and I woke up to this full prompt:
1
22
2,473
Mithil Vakde retweeted
Some cold takes… - I feel OpenAI is too modest in their defense. - Work of Córdoba and Martínez-Zoroa are fully public on Arxiv. Any agent swarm worth their salt will obviously look at these public results and would eventually explore those directions given enough compute. - It seems we still live in the world where we think humans are the center of the universe and only they have monopoly on insights. - Buckmaster et al obtained their headline results just past month. Given these models take >2 months to train, it seems extremely unlikely that OpenAI model might have seen their recent work. - Biggest takeaway for me is that compute scaling will keep working for at least 4 more OOMs. We have a perennial question when and if this scaling party will end. Seems we are safe (or unsafe in other ways) for 2 more years at least. - It’s cope when people try to undermine this achievement by saying they just threw huge compute. When humans work on something for 100 years, they are also just throwing more compute. - Another cope is people saying proof is slop because humans cannot understand it. Mathematics has no obligation for humans to understand it. Our intelligence is bounded and there will be longer time to digest things. Future prompt: “explain to me like I am human field medalist”. - Solving millennium prize problem is a massive milestone and we should celebrate this without any reservations. The reach of AI (and therefore humanity) has extended from solving 20 years old problems to almost 100 years. This just happened. - A sober take is that this achievement likely isn’t going to have parallels in physics/bio. Most of the wins here is verifiability and ton of RL for math. Other fields like physics/bio requires interaction with physical world for experimental verification and we might not see same acceleration there.
73
92
899
67,677
The most amazing thing I've seen. Can't wait for them to release all their results!
Replying to @OpenAI
This model represents a step-function improvement on many benchmarks, and its training is ongoing. Our internal model group arrived at the Navier–Stokes solution in 88 hours, using around 10,000 coordinating AI agents. Throughout the effort, we maintained the strict safeguards—including monitoring and isolation—that we apply to all our frontier evaluations.
2
12
1,138
Mithil Vakde retweeted
I would like to clarify a few things: 1) The screenshot is my reaching out to Levent to coordinate our releases. I hope it’s clear from the message that we came in with the best possible intentions. 2) I never ever asked for Levent to be removed from authorship of his own work (as indicated by my text). I was surprised to learn during the call with Tristan that they had only solved Euler and not Navier-Stokes; after learning this we brainstormed possible paths forward. One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work. Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic. Another option I wanted to propose (but got cut short) is to offer access to our internal model so that they could try to finish their proof and bridge the gap between Euler and NS. Again I did not know how to navigate giving access to internal OpenAI IP to an Anthropic employee. 3) To reiterate it plainly: as my text clearly indicates, and as I said during our call, OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler. In the call I was immediately met with a litany of slander, including direct threats that if we were to announce Navier-Stokes he would immediately go to the press with a barrage of unfounded accusations. I refuted all these accusations but he replied “there is nothing you can do, I simply do not trust you”. I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering, which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him and do a last ditch attempt to get a chance to give them all the credits that they deserve. I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey. (I should say that I retracted them on the spot by the way.) 4) Overall, on a personal level, it was incredibly difficult to have these conversations. Levent refused to attend any of the meetings despite my repeated asking. As Sholto Douglas said, there will need to be coordination between Anthropic and OpenAI in the future; I felt I was doing a proxy negotiation with Anthropic while the Anthropic employee refused to directly participate.
466
492
5,708
3,861,828
Mithil Vakde retweeted
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
5,704
20,138
120,423
75,086,811
Bubble? Bro, its a blowup!
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
2
1
55
1,649