We've figured out how to *pretrain transformers* with zeroth-order optimization and no backprop. Many of the core assumptions in optimization research are completely wrong. (paper out soon)

Sep 28, 2026 · 8:07 PM UTC

168
152
2,688
492,259
Full paper:
Backprop has been the only credit assignment algorithm capable of training large neural nets. Introducing Dust, the first zeroth-order method to pretrain transformers to approach and sometimes even *exceed* backprop with large amounts of computation (bitter lesson!) - Dust is on the order of 1,000 to 10,000x more compute efficient than EGGROLL, the state-of-the-art ES method, for training transformers. - We chose pretraining transformers because it's arguably the hardest task possible for zeroth-order optimization. - Search in high dimensions is really misunderstood: zeroth-order usually gets better with overparameterization, not worse, i.e. at a fixed population size, larger models (up to 120x) reach a lower loss than smaller ones. - Backprop-like gradients emerge from a large population of activation perturbations, and the alignment holds up at every scale we tested, up to 1B tokens, which is critical for scaling. The core idea is a *virtual population*, which helps scale to large populations in transformers by perturbing activations in parallel. w/ @bishmdl76, @cs_serdar, @akshayvegesna
9
801
Sort replies: Relevant Recent Liked
Replying to @industriaalist
holy hecking heck?!?!
2
53
15,751
lol
14
13,714
Replying to @industriaalist
Is this eggroll?
1
11
11,283
Replying to @industriaalist
nice! this is on text datasets?
1
8
13,676
yes fineweb
33
12,439
Replying to @industriaalist
is it meaningfully different from EGGROLL? arxiv.org/pdf/2511.16652
2
1
65
20,510
yes a lot
1
1
66
28,822
Replying to @industriaalist
what is the optimizer you feed the back-propogated gradients into (dashed line) ?
1
24
16,009
similar dynamics with all optimizers
2
26
15,106
Replying to @industriaalist
"Many of the core assumptions in optimization research are completely wrong." Utter nonsense. Optimization people understand zeroth-order optimization very well.
1
61
2,544
excited to read this
1
40
5,187
Replying to @industriaalist
Note how compute isn't on x axis...
31
1,363
Replying to @industriaalist
would be awesome if this works! convenient place to sink FLOPs in a token-limited world but it looks like the plot is actually cut off before backprop would have won? and it doesn't look like the curve had to end there because there's no dip from LR annealing at the end?
1
30
5,222
Replying to @industriaalist
evolution strategies?
2
21
5,813
Replying to @industriaalist
tyranny of backprop
1
21
2,961
Replying to @industriaalist
Idk what that means but it sounds like a big deal.
1
20
2,673
Replying to @industriaalist
the next year is going to be so bright.
14
2,262
Replying to @industriaalist
very impressive! i really need to finish writing my ZO blog it seems
2
13
3,489
Replying to @industriaalist
@grok what is a zeroth order optimization method? Something like a particle swarm?
1
4
1,542
Replying to @industriaalist
here for it
2
1,685
Replying to @industriaalist
Ohhh. So my heterogenous distributed training POC (I call the SETI@Home of training) will get an actual scalable algorithm to play with. Dying to see the paper.
2
818
Replying to @industriaalist
Perturbation?
1
683
Can’t wait 👀
1
1,529
Replying to @industriaalist
intriguing... how? :D
802
Replying to @industriaalist
Is this based on some variant of cma-es?
1
727
Replying to @industriaalist
👀 excited to see the paper !
1
611
Replying to @industriaalist
Looking forward to this!
241