Backprop has been the only credit assignment algorithm capable of training large neural nets.
Introducing Dust, the first zeroth-order method to pretrain transformers to approach and sometimes even *exceed* backprop with large amounts of computation (bitter lesson!)
- Dust is on the order of 1,000 to 10,000x more compute efficient than EGGROLL, the state-of-the-art ES method, for training transformers.
- We chose pretraining transformers because it's arguably the hardest task possible for zeroth-order optimization.
- Search in high dimensions is really misunderstood: zeroth-order usually gets better with overparameterization, not worse, i.e. at a fixed population size, larger models (up to 120x) reach a lower loss than smaller ones.
- Backprop-like gradients emerge from a large population of activation perturbations, and the alignment holds up at every scale we tested, up to 1B tokens, which is critical for scaling.
The core idea is a *virtual population*, which helps scale to large populations in transformers by perturbing activations in parallel.
w/
@bishmdl76,
@cs_serdar,
@akshayvegesna