We deleted loss.backward() and the transformer still learned. Our new research, Dust, is out.
Backprop has been the only credit assignment algorithm capable of training large neural nets. Introducing Dust, the first zeroth-order method to pretrain transformers to approach and sometimes even *exceed* backprop with large amounts of computation (bitter lesson!) - Dust is on the order of 1,000 to 10,000x more compute efficient than EGGROLL, the state-of-the-art ES method, for training transformers. - We chose pretraining transformers because it's arguably the hardest task possible for zeroth-order optimization. - Search in high dimensions is really misunderstood: zeroth-order usually gets better with overparameterization, not worse, i.e. at a fixed population size, larger models (up to 120x) reach a lower loss than smaller ones. - Backprop-like gradients emerge from a large population of activation perturbations, and the alignment holds up at every scale we tested, up to 1B tokens, which is critical for scaling. The core idea is a *virtual population*, which helps scale to large populations in transformers by perturbing activations in parallel. w/ @bishmdl76, @cs_serdar, @akshayvegesna
1
12
772
Bishwas retweeted
We've figured out how to *pretrain transformers* with zeroth-order optimization and no backprop. Many of the core assumptions in optimization research are completely wrong. (paper out soon)
168
152
2,688
492,558
Bishwas retweeted
We think computational depth is the missing scaling axis, i.e. we should be doing a lot more deep learning! Every other axis has been scaled by OOMs over the past few years (params, data, sparsity, test-time reasoning), but depth has been stuck at ~100 layers since GPT-3. We've found that LLMs are both *severely* depth-bottlenecked and bad at using the depth they have, and that architectural interventions that lift this bottleneck efficiently lead to gains that increase with compute. w/ @akshayvegesna
37
95
1,201
213,348
Bishwas retweeted
We've been obsessed with bending the scaling laws recently and our new paper shows that looping with model growth improves the scaling exponent, leading to gain that compounds with compute! Everyone assumes architectural changes only give constant factor gains and pretraining progress mostly comes from data (e.g. @dwarkesh_sp's recent post). We found that model growth, looping, and boundary operators result in compute multipliers over standard transformers that grow exponentially with each OOM of compute. - 1.55x at 1e20 FLOPs and 2.7x projected at 1e25. - Matches GPT-3 13B on CORE with 20x less compute w/ @charllechen, @akshayvegesna, @andrewgwils 🧵
12
49
446
75,339
excited!
a veryyy deep paper coming out soon
7
600
how much of this could be autoresearch looking at the progress from early '26 to now
6
226
Bishwas retweeted
the 1–3% gain is because their search algorithm is probably still quite primitive and inefficient. we've *already* had much bigger gains in pretraining. people think this is crazy, but searching over models is the core primitive missing from gradient descent, and adding that will improve generalization drastically. qlabs.sh/research
6
3
55
6,062
been 7 years and we are still doing this... attention is all you need (2017) search is all you need (2027?)
1
10
379
be on the lookout 👀
we'll replace gradient descent within the next 6 months. it'll be the biggest change in how we do deep learning possibly ever, bigger than rnns to transformers.
13
776
sounds crazy, but that's what we are seeing...
from everything i've seen, the answer to RSI is somewhere between 100T to 1000T parameters. big labs should go back to 2023 param-maxxing era
1
9
504
Really enjoyed listening to @SuryaGanguli talk about scaling laws and generalization. The discussion around how scaling behavior emerges from the structure of natural language and what that means for generalization was pretty fun. Sharing a short clip below:
1
1
10
323
looking forward to this!
this thursday we'll have jesse hoogland (@jesse_hoogland) to discuss singular learning theory (SLT). SLT establishes a connection between the geometry of the loss landscape and internal structure in models, offering a principled answer to why neural networks generalize. arxiv.org/pdf/2510.12077
7
256
Bishwas retweeted
this thursday we'll have surya ganguli (@SuryaGanguli) to discuss the origins of scaling laws. his paper predicts scaling exponents of LLMs from measurable statistics of natural language. more generally, his work draws on ideas from statistical physics to understand deep learning. personally, i think scaling laws are an incredible phenomenon that needs serious study, and understanding their origins is crucial for improving them. arxiv.org/abs/2602.07488
10
20
159
26,555
Bishwas retweeted
we'll have @AlexiGlad this thursday to discuss energy-based transformers. energy-based models (EBMs) are a more general framework for generative modeling than autoregression and diffusion (which is an implicit EBM). EBTs learn an explicit energy function for transformers, which allows test-time optimization over a learned landscape for stronger OOD generalization. register: luma.com/34g58cin arxiv.org/abs/2507.02092
3
30
9,053
we recently published a paper on how to make the most of a multi-epoch budget when pretraining language models on limited data. here's a blog post about it: bishwasmandal.com/blog/q0-hy…
1
7
20
1,942
Bishwas retweeted
this thursday evening we'll have andrew gordon wilson (@andrewgwils) over to discuss epiplexity, compression, and understanding generalization in LLMs. it'll be a small group, in-person, hosted at our office in sf. dm me to join.
2
3
51
11,294
Bishwas retweeted
hey everyone, i'm starting a tiny in-person reading group in SF for people working on generalization / science of deep learning. fundamentals matter more than ever and i think this'll become the biggest research area soon. we'll discuss big new ideas and bring in the best generalization researchers for talks. dm me if interested.
12
7
125
10,723