Research fellow at Flatiron Institute, working on understanding optimization in deep learning. Previously: PhD in machine learning at Carnegie Mellon.

New York, NY
Replying to @deepcohen
Part 1: How does gradient descent work? centralflows.github.io/part1… Part 2: A simple adaptive optimizer centralflows.github.io/part2… Part 3: How does RMSProp work? centralflows.github.io/part3…
1
11
124
18,195
Jeremy Cohen retweeted
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. 1/N
43
187
1,413
279,091
Jeremy Cohen retweeted
We've been working for about a year now on how to port the core idea of Muon to LoRA (low rank) finetuning. If you're interested in Muon, preconditioned Muon, Lora, let me tell you about what we learned and our PoLoRA method arxiv.org/abs/2607.17620 nikhilgsh.github.io/polora/
6
27
164
13,947
Did Anthropic get more gains out of model scaling than other labs thought was possible? It reminds me of an interesting recent paper, which showed that deep layers in open LLMs are not doing much, and that this can be fixed by scaling the LayerNorm output. arxiv.org/abs/2502.05795
9
22
298
25,810
Update: not the secret sauce (definitely known to the other labs)
Replying to @deepcohen
Thanks for your kind thoughts and support, Jeremy! I do believe this direction could potentially open up a new level of scaling toward deeper models with stronger reasoning capabilities. To be clear, I was not aware of the Depth-MP work when developing this paper (shamed); the idea grew out of my years of work on model compression and layerwise analysis. 1. arxiv.org/abs/2202.02643 2. arxiv.org/abs/2310.05175 3. arxiv.org/abs/2410.10912 After releasing our paper, we compared our method with Depth-P and found that the two perform similarly. I fully respect prior work, and we will expand the related work discussion and add the appropriate references in the next version.
1
7
1,378
Would be interested in hearing others’ wild speculation
741
The recent Microsoft AI report noted that too much learning rate decay during pretraining hurts post-RL performance. This is actually just the latest of several papers this year pointing out that small learning rates can be harmful in LLM pretraining. (Thread)
5
23
198
18,550
Nevertheless, hyperparameters matter, and I'm glad we're starting to see good science about LR schedules for LLMs. PS: as noted by Catalan-Tatjer et al, weight-averaging recovers many benefits of LR decay but without increasing sharpness. More should consider weight-averaging!
1
2
14
987
Oh also, in the *multi-epoch* LLM setting, there is evidence that larger LR's yield better population pretraining loss, exactly mirroring what was known in 'classical' image settings (arxiv.org/abs/2306.08590). (This experiment is with batch size, but LR should be the same)
1
5
640
Jeremy Cohen retweeted
1/ Deep learning is going to have a scientific theory. We can see the pieces starting to come together, and it's looking a lot like physics! We're releasing a paper pulling together these emerging threads and giving them a name: learning mechanics. 🔨 arxiv.org/pdf/2604.21691 🔧
51
285
1,495
311,587