Prof at @GatsbyUCL and @SWC_Neuro, trying to figure out how we learn. Bluesky: @SaxeLab Mastodon: @SaxeLab@sigmoid.social

London, UK
Filter
Exclude
Time range
-
Minimum likes
Come chat about this @iclr_conf, at 3:15 PM on Friday in Pavilion 4 Poster #4216!
Why don’t neural networks learn all at once, but instead progress from simple to complex solutions? And what does “simple” even mean across different neural network architectures? Sharing our new paper @iclr_conf led by Yedi Zhang with Peter Latham arxiv.org/abs/2512.20607
6
51
6,013
Very excited by this year's Analytical Connectionism Summer School! A dream lineup of speakers on the topic of language acquisition in minds and machines Bursaries available to cover costs Aug 17 – Aug 28, 2026 Gothenburg Details: analytical-connectionism.net…
6
33
3,367
We’re hiring postdocs/research scientists! Your interests can be anywhere on the spectrum from pure theory to empirically testing predictions relevant to AI safety. Our theoretical work relies on dynamical systems and tools from statistical physics. 3
2
2
47
3,024
We avoid many dangerous outcomes in the physical world using our knowledge of physics, and basic deep learning theory should eventually enable the same for AI. We focus on analytically tractable “model organisms” that capture essential learning dynamics and failure modes. 2
1
29
2,255
Excited to launch Principia, a nonprofit research organisation at the intersection of deep learning theory and AI safety. Our goal is to develop theory for modern machine learning that can help us understand network behaviors, including those critical for AI safety. 1
9
35
301
19,316
Equipped with this theory, we make new predictions about how network width, data distribution, and initialization affect learning dynamics. For example, increasing the number of attention heads in linear attention shortens the plateaus in learning.
2
9
841
So when progressing simple -> complex, linear networks learn solutions of increasing rank, ReLU networks learn solutions with increasing kinks, convolutional networks learn solutions with increasing convolutional kernels, and attention models learn solutions with increasing heads
1
1
10
697
Here the notion of simplicity is the number of effective units in the architecture: hidden neurons, convolutional kernels, or attention heads.
1
6
540
Finally, we demonstrate gradient descent sometimes naturally evolves along the connecting paths between saddles iteratively, yielding saddle-to-saddle dynamics. We identify two distinct mechanisms: timescale separation between directions or units, depending on the architecture.
1
1
7
595
We then show that saddles are connected by gradient descent paths (invariant manifolds). Along these paths, a larger network behaves like a smaller one, retaining the same simplicity during a saddle-to-saddle transition.
1
2
6
844
We first show that saddle points are ubiquitous in the loss landscape: fixed points of smaller networks can be embedded as saddle points of larger networks, yielding a nested hierarchy of saddles. These saddles exist in any network that contains a sum of repeated units.
1
3
10
1,250
We present a theoretical framework that explains a dynamical simplicity bias arising from saddle-to-saddle learning dynamics across neural network architectures: Fully-connected, convolutional, attention-based, and more. yedizhang.github.io/simplici…
1
1
15
1,228
Why don’t neural networks learn all at once, but instead progress from simple to complex solutions? And what does “simple” even mean across different neural network architectures? Sharing our new paper @iclr_conf led by Yedi Zhang with Peter Latham arxiv.org/abs/2512.20607
11
59
440
30,356
Come chat about this at the poster @icmlconf, 11:00-13:30 on Wednesday in the West Exhibition Hall #W-902!
How does in-context learning emerge in attention models during gradient descent training? Sharing our new Spotlight paper @icmlconf: Training Dynamics of In-Context Learning in Linear Attention arxiv.org/abs/2501.16265 Led by Yedi Zhang with @Aaditya6284 and Peter Latham
4
15
2,190
With this work we hope to shed light on some of the universal aspects of representation learning in neural networks, and one way networks can generalize infinitely from finite experience. (11/11)
8
522
Not all pairs of representations which can merge are actually expected to merge. Thus, the final learned automaton may look different for different training runs, even when in practice they implement an identical algorithm. (10/11)
2
6
567
The theory predicts mergers can only occur given enough training data and small enough initial weights, resulting in a phase transition between an overfitting regime and an algorithm-learning regime. (9/11)
1
5
183