The Center for Brains, Minds and Machines is a multi-institutional NSF Center dedicated to the study of the science and engineering of intelligence.

Cambridge, MA
Edge of Stability has been a thing for a while now. We can show that most training happens at the edge of stability. This is very surprising because of a "negative " fact: Most theory does not apply (e.g., Descent Lemma). The question becomes: Does the fact that training happens at the EoS impact performance? In what way? With what mechanisms? How is that affected by the architecture? To our knowledge, these questions they are unanswered to date and our new paper is the about the first step towards this! EoS picks what part of the distribution to learn and at what speed! arxiv.org/abs/2606.04212 In particular, the consequence is that for MLPs: EoS can improve robustness or OOD behavior, but only when the relevant subset is the one selected by EoS. If boundary points dominate, EoS helps near-boundary robustness. If distributional outliers dominate, EoS helps extrapolate toward the tail. Precisely: 1. We make EoS causal. We fork training from the same state at EoS onset: one branch stays at EoS, the other exits by lowering the learning rate. Same everything, only the stability constraint changes. 2. EoS is selective: it learns some parts of the data distribution faster, while slowing others. 3. The selector is surprisingly simple. A group benefits from EoS when its gradient is big and has high alignment with v_1, the top Hessian eigenvector. Translation: to benefit from EoS, a subset must (i) point in the sharpest direction, and (ii) keep a non-vanishing gradient. This gives concrete answers to the questions above: Does EoS affect performance? Yes, but not uniformly. It reallocates learning. In what way? It prioritizes the subset with largest curvature influence. By what mechanism? Self-stabilization at the sharpness boundary couples training to v_1; only groups aligned with v_1, and whose gradients persist, get the extra progress. How does architecture matter? Architecture changes the map from input geometry to gradient geometry. Thus it can change which subset dominates, but the same predictor remains: largest curvature influence wins. With the amazing Shauna Kwag, @anakha_g, and @TomasoPoggio!
1
12
95
6,706
CBMM retweeted
Excited to share our theory on the universal compressibility of neural nets (led by @HYWangphy)! This leads to a new dynamical lottery ticket hypothesis and implications for improving neural scaling laws. arxiv.org/abs/2510.00504 Check out our ICLR poster!
4
14
128
10,241
CBMM retweeted
This is a very cool idea to decomposes sequence modeling into learning what to remember and how to update memory!
We never really knew how to train nonlinear RNNs well… BPTT struggled with vanishing grads (no long-range memory) and sequential rollout (hard to parallelizable). What if instead an oracle told us the optimal memory state m_t at each step? Then the RNN could do one-step supervised learning on (m_t, x_{t+1}) → m_{t+1} labels. We call this Supervised Memory Training (SMT): a replacement for BPTT that trains RNNs without unrolling them. SMT is time-parallelizable and solves vanishing gradients. Website: akarshkumar.com/smt/ arXiv: arxiv.org/abs/2606.06479
1
3
18
4,415
A common practitioner belief mechanistically: “weight decay stabilizes training.” We show it is wrong technically, a bit real In practice. - Training still happens at EoS with weight decay. - That said "it doesn't look like it"! The curvature at stabilization is much lower! - We properly charcaterize all we see by extending to weight decay the self-stabilization by @alex_damian_, @EshaanNichani, and @jasondeanlee and matching the result cleanly with an underdamped harmonic oscillator! - This however stabilizes/tames the dynamics in function space. Importantly: - This shows that regularized training at EoS even though "doesn't look like that", meaning curvature measures are lower than thresholds. Check out our paper: arxiv.org/pdf/2605.16622
2
11
87
6,050
Check out the latest work of our center, in collaboration with @TAMU! Towards theorizing the boost in capabilities of agent systems. @PierBeneventano @GalantiTomer
Have you ever wondered how to formalize what an agentic system actually is? Meaning where they fit in the book of ML and how to explain/predict their performance? We argue here, agents can be seen as boosting reasoning models! arxiv.org/abs/2605.14163
2
6
914
1/ Many optimization problems are hard in theory. But real OR and NP-hard instances often have exploitable structure. Can an LLM agent discover that structure automatically and turn it into faster solver code?
5
34
205
25,756
[blog] What is Intelligence? Or "Distinguishability is All You Need" Here are several related questions to which we do not have a good answer: How will we know when we've achieved "Artificial General Intelligence" (AGI)?... poggio-lab.mit.edu/blogsupda…
235
[video] "Intelligence as Prediction: Cybernetics, LLMs, and Sociality" Speaker: Blaise Agüera y Arcas - Google, Paradigms of Intelligence piped.video/6NC0tSjZXBo
3
16
1,218
[blog post] "PoggioAI/MSc Went Online" This first public release is an open-source, customizable, modular multi-agent system for academic research workflows, with a current emphasis on machine learning theory and nearby quantitative fields. poggio-lab.mit.edu/blogsupda…
1
2
475
Check the blog of Poggio Lab at MIT! We went online with some very nice blogs! The last one being about our multiagent system: poggio-lab.mit.edu/blogsupda…
3
6
956
Most AI for research work tries to maximize autonomy first and patch quality later. We think the near-term path is the reverse: Automating step-by-step holding the quality bar fixed. Today we’re open-sourcing PoggioAI/MSc for ML Theory Research
1
8
34
43,398
CBMM retweeted
Simply adding Gaussian noise to LLMs (one step—no iterations, no learning rate, no gradients) and ensembling them can achieve performance comparable to or even better than standard GRPO/PPO on math reasoning, coding, writing, and chemistry tasks. We call this algorithm RandOpt. To verify that this is not limited to specific models, we tested it on Qwen, Llama, OLMo3, and VLMs. What's behind this? We find that in the Gaussian search neighborhood around pretrained LLMs, diverse task experts are densely distributed — a regime we term Neural Thickets. Paper: arxiv.org/pdf/2603.12228 Code: github.com/sunrainyg/RandOpt Website: thickets.mit.edu
91
462
3,170
789,883
[blog] Beneficial Misalignment: Why We Shouldn't Always Align AI to Humans In the rapidly evolving field of NeuroAI, a significant amount of energy is dedicated to 'alignment', the idea that representations from artificial intelligence should converge... poggio-lab.mit.edu/blogsupda…
1
10
704
[blog post] A Conversation with Blaise Agüera y Arcas: On Intelligence, Life, and the Future of AI What does it mean to call something intelligent - and when did this question get so hard to answer? For Blaise Agüera y Arcas, VP at Google and founder... poggio-lab.mit.edu/blogsupda…
1
1
6
964
[blog post] Can a Neural Network Think Before It Speaks? Somewhere around 2022, an observation started making the rounds among researchers working with large language models: if you just asked a model... poggio-lab.mit.edu/blogsupda…
7
638
[blog post] Edge of (Stochastic) Stability made simple — Part II: the mini-batch case In Part I we had one landscape and a deterministic update. Now we have a distribution of mini-batch landscapes and a stochastic update... poggio-lab.mit.edu/blogsupda…
1
299
[blog post] Edge of (Stochastic) Stability made simple — Part I: A crash course on (full-batch) Edge of Stability In this part I introduce the phenomenon and what I believe are the two key mechanisms—which we’ll use as the springboard for the mini-bat... poggio-lab.mit.edu/blogsupda…
2
7
607
[blog post] Are Transformers Just "Stochastic Parrots"? A common criticism of Large Language Models (LLMs) is that they are merely "stochastic parrots"—statistical mimics that stitch together likely patterns without genuine reasoning... poggio-lab.mit.edu/blogsupda…
1
6
420
🧵 New paper: LLM-ERM: Sample-Efficient Program Learning via LLM-Guided Search arxiv.org/abs/2510.14331 We use reasoning LLMs to learn tasks like IsPrime from ~200 samples by proposing short programs, making both the learned function *and* the learning process interpretable 🤯
2
10
36
8,558
Does SGD really “seek flat minima”? We show that SGD has no intrinsic preference for flatness, even for stable linear networks—going against ~10 years of folklore. Flatness emerges iff label noise is isotropic; anisotropic noise drives SGD to arbitrarily sharp solutions. This reveals a new flattening–sharpening mechanism in late training, unrelated to standard progressive sharpening or Edge-of-Stability effects.
10
36
395
24,751