Our paper, “Make Haste Slowly: A Theory of Emergent Structured Mixed Selectivity in Feature Learning ReLU Networks” will be presented at ICLR 2025 this week (openreview.net/forum?id=27SS…)! We derive closed-form dynamics for some, remarkably linear, feature learning ReLU networks (1/9)
Apr 21, 2025 · 2:20 PM UTC
1
15
67
4,878
Our main idea is to map ReLU networks onto Gated Deep Linear Networks (GDLNs) (Saxe et al., 2022). GDLNs are neural networks which nonlinearly combine linear neural networks using gating operations. Importantly, GDLNs permit a dynamics reduction for their learning dynamics. (2/9)
1
1
144
We define the Rectified Linear Network (ReLN) as the GDLN which has the same output as the ReLU network at all points in time and prove that a ReLN always exists for a given ReLU network. Thus, we obtain the dynamics of learning for the ReLU networks as well! (3/9)
1
108
We demonstrate the paradigm on an extended XoR task which includes a third feature, making the dataset linearly separable. We observe a transition from the ReLU network using the linear strategy to the nonlinear strategy based on the magnitude of the new feature. (4/9)
1
65
We then consider a complex nonlinear, contextual task. We find that the ReLU network has an inductive bias towards mixed-selective latent representations where no hidden neuron is selective for an item or context. Instead it couples linear pathways to favour learning speed. (5/9)
1
65
We obtain closed-form dynamics for the ReLU network in this setting and see near perfect agreement with the predicted dynamics. We prove that the identified ReLN is unique and show that the mixed-selectivity preference remains when the number of contexts is increased. (6/9)
1
61
We consider the effect of adding depth to the ReLU network and find that the network still favours linear pathways with mixed-selective representations. However, there is variance in the dynamics and we show a corresponding GDLN can still model the distribution of dynamics. (7/9)
1
65
Finally, we provide an initial hidden layer clustering algorithm which is able to identify a ReLN for a given ReLU network, with the goal of enabling future work with our paradigm. (8/9)
1
123
A huge thanks to my supervisors @kleinric, @BenjaminRosman and @SaxeLab for their constant assistance and guidance. Looking forward to chatting more on Thursday at 3pm, Poster #349! 👋 (9/9)
1
113







