We've figured out how to *pretrain transformers* with zeroth-order optimization and no backprop.
Many of the core assumptions in optimization research are completely wrong.
(paper out soon)
Sep 28, 2026 · 8:07 PM UTC
168
152
2,688
492,259
Full paper:
Backprop has been the only credit assignment algorithm capable of training large neural nets.
Introducing Dust, the first zeroth-order method to pretrain transformers to approach and sometimes even *exceed* backprop with large amounts of computation (bitter lesson!)
- Dust is on the order of 1,000 to 10,000x more compute efficient than EGGROLL, the state-of-the-art ES method, for training transformers.
- We chose pretraining transformers because it's arguably the hardest task possible for zeroth-order optimization.
- Search in high dimensions is really misunderstood: zeroth-order usually gets better with overparameterization, not worse, i.e. at a fixed population size, larger models (up to 120x) reach a lower loss than smaller ones.
- Backprop-like gradients emerge from a large population of activation perturbations, and the alignment holds up at every scale we tested, up to 1B tokens, which is critical for scaling.
The core idea is a *virtual population*, which helps scale to large populations in transformers by perturbing activations in parallel.
w/ @bishmdl76, @cs_serdar, @akshayvegesna
9
801



























