two Lagrangian papers published in 2026 next year is year of the goat i.e. back to our regularly scheduled hill climbing
2026: year of the horse i.e. year of the saddle point. Brushing up on Lagrange duality for a head start.
2
16
1,097
I’ve purchased and subsequently lost 87 umbrellas in the past 36 days
🇯🇵 Tokyo has recorded rain for 36 consecutive days, the longest streak since records began in 1876. Aug 27: 🌧️ Aug 28: 🌧️ Aug 29: 🌧️ Aug 30: 🌧️ Aug 31: 🌧️ Sep 1: 🌧️ Sep 2: 🌧️ Sep 3: 🌧️ Sep 4: 🌧️ Sep 5: 🌧️ Sep 6: 🌧️ Sep 7: 🌧️ Sep 8: 🌧️ Sep 9: 🌧️ Sep 10: 🌧️ Sep 11: 🌧️ Sep 12: 🌧️ Sep 13: 🌧️ Sep 14: 🌧️ Sep 15: 🌧️ Sep 16: 🌧️ Sep 17: 🌧️ Sep 18: 🌧️ Sep 19: 🌧️ Sep 20: 🌧️ Sep 21: 🌧️ Sep 22: 🌧️ Sep 23: 🌧️ Sep 24: 🌧️ Sep 25: 🌧️ Sep 26: 🌧️ Sep 27: 🌧️ Sep 28: 🌧️ Sep 29: 🌧️ Sep 30: 🌧️ Oct 1: 🌧️
1
21
1,343
Next year's NeurIPS registration site
6
397
Accepted to NeurIPS! 🇦🇺
sharing a new paper! Augmented Lagrangian Predictive Coding I think one of the coolest unsolved problems is how the brain does credit assignment. thread 🧵
8
3
82
5,470
Jeffrey Seely retweeted
Introducing the Sakana AI Frontier Intelligence Group 🪷 sakana.ai/frontier-intellige… Current AI systems are incredibly capable, but is intelligence “solved”? And if not, what’s missing? At Sakana AI’s Frontier Intelligence Group (FIG), we believe that there are still breakthroughs to be made in AI. The Transformer and language modeling may be incredibly powerful, but it doesn’t mean that better alternatives don’t exist. Natural intelligence still beats artificial intelligence across many dimensions. Agents lack the deep insights and creativity of humans. Individual models require far more data than the brain to learn robustly, and require far more energy to run. If we set these as targets, what kinds of AI systems could we develop? Research at FIG has sought to address the gaps between natural and artificial intelligence. Here are some of our works, and the fundamental research questions that motivated them: • Continuous Thought Machines: How can we improve information processing by leveraging temporal dynamics? • Augmented Lagrangian Predictive Coding: How can local learning solve multilayer credit assignment? • Sparser, Faster, Lighter Transformer Models: How can we massively increase data efficiency and generalization? • The AI Picbreeder Experiment: How can we make artificial open-ended systems? • Smart Cellular Bricks: How can physical systems achieve collective intelligence and self-repair without a central brain? We hope that this encourages other researchers to also explore different paradigms, and take a leap of faith with us. After all, in the words of a dear friend of ours, “greatness cannot be planned”.
37
87
772
131,948
Jeffrey Seely retweeted
Replying to @SakanaAILabs
Nice work! I was immediately reminded of the sheaf-theoretic implications here and then I found @jeffreyseely's paper about it (arxiv.org/abs/2511.11092) Using λi to integrate the sheaf coboundary residual is a neat way to get the PI controller behavior. I am curious if you ran any experiments using a local rate to calculate a non-plant derivative term for the controller to accelerate the gluing dynamics? Something like vi = beta * vi + (1 - beta)(ri - r[i-1])?
1
2
5
1,343
Jeffrey Seely retweeted
I’ve been tracking this work closely since I met @jeffreyseely at ICML, and it’s awesome to see this go live. I think there’s something so cool about that wave-like visual he produced showing how you can make predictive coding distribute its error signal about as well as backprop.
sharing a new paper! Augmented Lagrangian Predictive Coding I think one of the coolest unsolved problems is how the brain does credit assignment. thread 🧵
1
1
9
599
Jeffrey Seely retweeted
Replying to @SakanaAILabs
Nice! This ends up being a version of what some of us have called "target prop": every layer's input is a free latent variable that serves as a target for the previous layer. As this paper points out, this can be derived from an "augmented Lagrangian" formulation of backprop in which the constraints (input of layer k+1 = output of layer k) are turned into penalties (divergence between input of layer k+1 and output of layer k). I've always hoped more people would pick up on this idea. I'm happy this is happening! I must say though that target prop, in the end, optimizes the same criterion as backprop and does the same thing as backprop while evaluating the gradient in a different way, perhaps more biologically plausible. My lab did some work on this idea in the context of "sparse auto-encoders" in the late 2000s. It turns out when the code in an auto-encoder is regularized (e.g. with L1 to make it sparse) target prop seems more efficient than backprop. scholar.google.com/citations…
19
54
621
50,690
sharing a new paper! Augmented Lagrangian Predictive Coding I think one of the coolest unsolved problems is how the brain does credit assignment. thread 🧵
Introducing PC-ALM, a local-learning alternative to backpropagation. Our method trains 1000-layer neural nets using only local dynamics, and without backprop. Blog: pub.sakana.ai/pc-alm/ Standard deep learning relies on backpropagation. The brain, however, cannot implement backpropagation, at least not exactly. How can a physical system, such as the brain, solve multilayer credit assignment without explicit use of backprop? We look for inspiration in two related fields: distributed optimization and NeuroAI. In NeuroAI, predictive coding asks each neuron activation to solve an energy-based inference problem instead of using a standard forward pass. That inference step can be implemented as energy-minimization dynamics on local prediction errors. This perspective -- each layer as a dynamical system -- has proven promising, but performance of predictive coding hasn't scaled well with depth. Credit signals at far ends of the network struggle to diffuse into internal layers. We turn to distributed optimization, generalizing predictive coding to use an augmented Lagrangian instead of energy. This motivation stems back to a classic 1988 paper by LeCun, showing that the Lagrange multipliers of a deep network can be identified with gradients of a supervised loss. The augmented Lagrangian then bridges LeCun's perspective to the standard predictive coding that is used in NeuroAI. We find that this new perspective yields a natural PC-like alternative to backpropagation, resulting in a method we call PC-ALM. PC-ALM differs from PC in that it introduces dual neurons (Lagrange multipliers) as part of the layer-local dynamics, resulting in each layer acting as a PI feedback control system to minimize local prediction errors. We find that PC-ALM is capable of propagating signals to seemingly arbitrary depth, especially in deep narrow networks where standard PC struggles to learn. Ultimately, our motivation here is to understand how distributed physical systems, such as the brain, can compute credit signals using only local coupling and local dynamics. PC-ALM may also inform deep learning in neuromorphic hardware, where dynamics are cheaper than on GPUs. Paper: arxiv.org/abs/2605.31022 Code: github.com/SakanaAI/pc-alm
12
41
315
33,908
How to improve PC then? We look for clues in distributed optimization. We use the augmented Lagrangian instead of the PC energy. We do primal-descent dual-ascent on the AL (which we dub PC-ALM). Surprisingly, this simple modification drastically improves credit propagation in deep narrow networks compared to PC, while also introducing more realistic neural response profiles.
1
14
764
I particularly like the perspective of primal-descent dual-ascent as a (PI) feedback control system. Global credit assignment emerges from a network of local feedback controllers! See the paper and our blog (which I spent too much time on) and our code! blog: pub.sakana.ai/pc-alm/
2
1
27
812
Maybe one outcome of this is that more and more AI people will read William Thurston's On Proof and Progress in Mathematics, which is just great.
1
1
14
952
The perfect tweet
Replying to @NGKabra
That's why they call it THE SINGULARITY
10
3,216
Anyone else wonder why Clay institute didn’t dump the funds into S&P500
3
23
4,598
Jeffrey Seely retweeted
If you are curious about "Active Inference" and how brains might plan actions using hierarchical active inference, I invite you to read this new paper by my PhD student Prashant Rangarajan, published recently in the journal Neural Computation. @uwcse @UW direct.mit.edu/neco/article/…
13
44
3,297
Norwegian wedding music
1
13
1,677