Calculation Consulting is a boutique consultancy that specializes in machine learning, AI, and data science

San Francisco
Calc Consulting retweeted
If you train too long, you can actually collapse generalization The authors call this anti-grokking Typically what you see in grokking is training accuracy saturating while test accuracy remains low, followed by a surge where the model finally generalizes. What happens if you keep training beyond that peak? test accuracy can actually suddenly collapses back to chance levels, even as training accuracy remains perfect. How can we detect this? Standard grokking progress measures (L2​ weight norms, activation sparsity, weight entropy) track the initial grokking phase but stay entirely flat and miss the "anti-grokking" Instead, using the WeightWatcher diagnostic tool, the authors analyze the empirical spectral density of the layer weight matrices. anti-grokking is driven by the emergence of "Correlation Traps"—aka weirdly large eigenvalues that appear late in training and severely impair generalization. The Heavy-Tailed Self-Regularization (HTSR) metric tracks this; At peak generalization, model layers reach an optimal alpha~2. On a 3-layer MLP (MNIST), alpha drops below 2, indicating "prototype overfitting" where the model memorizes specific training instances and gets confused. On a Transformer (Modular Addition), alpha climbs above 2, indicating catastrophic forgetting. these correlation traps are also observed in GPT OSS 20B/120B
4
16
146
5,442
While there is benign overfitting in LLMs, not all overfitting is benign.
5
387
I hate to say I told you so… no actually I don’
The hammer of justice has smashed @AnthropicAI arguments. They are a Supply Chain Risk to the defense industrial base serving the @DeptofWar. Warfighters will sleep better knowing that no private company will insert their opinions in the chain of command. @SecWar was right!
1
512
Calc Consulting retweeted
holy shit i asked claude to make a video on western civiization
2,655
10,114
58,239
12,612,746
Calc Consulting retweeted
Stanford dropped CS329Z today. Engineering AI Agents. compound systems, build RAG/tool use from scratch, then frameworks, then how you actually evaluate this stuff. gonna follow along online. looks pretty neat. cs329z.stanford.edu/
13
187
1,484
83,973
Here is a preliminary analysis of applying WeightWatcher to study how LLMs memorize random inputs. When we deliberately corrupt part of the training data with random labels, the model eventually memorizes those labels—and its clean test accuracy begins to deteriorate. We monitored the weight spectra throughout training and found a surprisingly strong, layer-specific signal. The clearest effect occurs in the attention K matrix. Across multiple independent seeds, the WeightWatcher power-law exponent α falls toward and then below the critical value α = 2 as random-label memorization becomes extensive. The relationship is remarkably reproducible, with a mean rank correlation of roughly ρ ≈ −0.90 between K α and memorization. The effect survives WeightWatcher’s fix_fingers correction, suggesting it is not simply an artifact of a few anomalous eigenvalues. Importantly, this is not a universal α < 2 rule Other layers behave quite differently. The emerging picture is that memorization involves a layer-selective spectral reorganization, with the K matrix providing an especially strong marker of the transition from generalization toward memorization and overtraining. These are still small NanoGPT experiments and the analysis is preliminary, but the multi-seed results are encouraging. The next step is to determine whether the same spectral signature survives at larger scale and across optimizers. If anyone is interested in getting involved or trying it themselves, please reach out here or on the weightwatcher ai community Discord
5
2
21
833
Calc Consulting retweeted
We can now do effective continual learning for #LLMs I am happy to present Info-SDFT: Information-Proximal Self distillation. The method nicely balances plasticity versus stability and works very well. We did over 500 experiments on this! Thanks to @KickItLikeShika! We open source code and checkpoints for you to use: Code: github.com/KickItLikeShika/i… Checkpoints: huggingface.co/collections/K… Paper: arxiv.org/pdf/2609.24646
5
14
194
12,034
Reinforcement Learning is the key to getting LLMs to appear to reason. If you know some statistical mechanics, you may, like I did, immediately recognize that the basic REINFORCE algorithm is essentially non-Equilibrium Monte Carlo, where the RL updates follow a gradient that looks exactly like a linear relaxation described the Fluctuation-Dissipation theorem. Here, I have vibed up a short document and explainer video to help explain the idea. I hope it is useful and is sparks some new ideas. Join us on the weightwatcher.ai discord to discuss
6
5
110
6,386
and someone on LinkedIn pointed out these recent references THERMODYNAMICS OF REINFORCEMENT LEARNING CURRICULA arxiv.org/pdf/2603.12324 Bridging Control, Inference, Transport, and Thermodynamics From Theory to Applications in Learning arxiv.org/pdf/2609.15897
1
4
295
things are getting weird
4
341
🔥Hot take: RL algos like REINFORCE and PPO are effectively non-equilibrium Monte Carlo. You perturb the system (prompt) and follow the gradient of the (generated) response along the linear response dictated by the fluctuation-dissipation theorem.
2
19
1,303
ICLR, here we come.
If you train a model for too long, it may overfit it's training data. Not surprising, this has been know for like forever. But did you know you can detect the signatures of overfitting in the layer weight matrices directly, without needing access to any data (train or test) ? In our recent paper (with hari kishan prakash ), 𝐋𝐚𝐭𝐞-𝐒𝐭𝐚𝐠𝐞 𝐆𝐞𝐧𝐞𝐫𝐚𝐥𝐢𝐳𝐚𝐭𝐢𝐨𝐧 𝐂𝐨𝐥𝐥𝐚𝐩𝐬𝐞 𝐢𝐧 𝐆𝐫𝐨𝐤𝐤𝐢𝐧𝐠: 𝐃𝐞𝐭𝐞𝐜𝐭𝐢𝐧𝐠 𝐀𝐧𝐭𝐢-𝐆𝐫𝐨𝐤𝐤𝐢𝐧𝐠 𝐰𝐢𝐭𝐡 𝐖𝐞𝐢𝐠𝐡𝐭𝐖𝐚𝐭𝐜𝐡𝐞𝐫, we show this explicitly in 2 different classic grokking experiments. And the overfitting we see is very different from what has been seen before! 📄 paper: arxiv.org/abs/2602.02859 📈 𝐖𝐡𝐚𝐭 𝐭𝐡𝐢𝐬 𝐩𝐥𝐨𝐭 𝐬𝐡𝐨𝐰𝐬: Grokking → Stability → Anti-Grokking This figure below tracks training accuracy (red), test accuracy (purple), and WeightWatcher correlation traps (blue) while training for very long times. 𝐏𝐡𝐚𝐬𝐞 𝟏 — Memorization (pre-grokking) Training accuracy rises rapidly while test accuracy remains low. The model is fitting the training data without extracting the underlying structure. Correlation traps are minimal and largely uninformative. 𝐏𝐡𝐚𝐬𝐞 𝟐 — Grokking Test accuracy suddenly jumps to match training accuracy. The model transitions from memorization to true generalization. Correlation traps remain near zero, indicating stable, well-conditioned internal representations. 𝐏𝐡𝐚𝐬𝐞 𝟑 — Late-stage instability (anti-grokking) Despite perfect training accuracy, test accuracy degrades over time. At the same time, correlation traps increase sharply and spread. Generalization collapses after it was achieved. The model is overfit 𝐊𝐞𝐲 𝐭𝐚𝐤𝐞𝐚𝐰𝐚𝐲 So what ? It turns out, a lot of open-source LLMs, like OpenAI's GPT OSS 20B and 120B , show the exact same signatures! And a ton of them If you are training or fine-tuning your own models, watch out! You might be overfitting your data, even if you are following current NN best practices. Want to learn more? Check out the WeightWatcher project 🌐 weightwatcher.ai 𝐖𝐞𝐢𝐠𝐡𝐭𝐖𝐚𝐭𝐜𝐡𝐞𝐫 𝐢𝐬 𝐚 𝐨𝐧𝐞-𝐨𝐟-𝐚-𝐤𝐢𝐧𝐝 𝐦𝐮𝐬𝐭-𝐡𝐚𝐯𝐞 𝐭𝐨𝐨𝐥 𝐟𝐨𝐫 𝐚𝐧𝐲𝐨𝐧𝐞 𝐭𝐫𝐚𝐢𝐧𝐢𝐧𝐠, 𝐝𝐞𝐩𝐥𝐨𝐲𝐢𝐧𝐠, 𝐨𝐫 𝐦𝐨𝐧𝐢𝐭𝐨𝐫𝐢𝐧𝐠 𝐃𝐞𝐞𝐩 𝐍𝐞𝐮𝐫𝐚𝐥 𝐍𝐞𝐭𝐰𝐨𝐫𝐤𝐬 (𝐃𝐍𝐍𝐬). And you need help with AI, reach out. hashtag#TalkToChuck P.S. I'll be giving a talk at USF right here in SF (over by the GG Park) in 2 weeks on the weightwatcher project. Hope to see you there!
2
7
145
29,240
ICLR Submission Number: 53755
10
8
161
44,786
𝗔𝗱𝗮𝗺𝗪 𝘃𝘀. 𝗠𝘂𝗼𝗻: α, 𝗢𝘃𝗲𝗿𝗳𝗶𝘁𝘁𝗶𝗻𝗴, 𝗮𝗻𝗱 𝗠𝗲𝗺𝗼𝗿𝗶𝘇𝗮𝘁𝗶𝗼𝗻. A quick update on our WeightWatcher experiments comparing AdamW with the newer Muon optimizer. We trained a single-head nanoGPT model with both optimizers across five seeds. We then used WeightWatcher to study the weight-matrix spectra and several forms of memorization. The top plot shows the average raw α. AdamW produces relatively stable α values. Muon is much noisier. Its ESDs vary substantially across layers, which makes α harder to estimate reliably. But the average hides something important: 👉 Individual layers do fall below α = 2. Our spectral/RG theory predicts that this regime can be associated with pathological, example-specific learning. So we tested that prediction directly. We inserted random sequences into the training set and measured whether the model remembered them using exact extraction, token recall, teacher-forced recall, and canary exposure. The result was striking. • AdamW memorized the planted sequences strongly. Earlier in training it reached ~96% exact recall of 32-token canaries. • Muon strongly suppressed this kind of memorization. At 10,000 steps, teacher-forced recall was ~62% for AdamW versus <1% for Muon. There is also a surprise: some Muon layers have α < 2 as well. So α < 2 by itself is not sufficient to explain memorization. This is why we recommend examining individual WeightWatcher layers and their ESDs—not simply average α. Muon appears to suppress memorization while also producing unusual ESDs that make conventional WeightWatcher analysis more difficult. We are now running Muon much longer to understand how these spectra evolve. If you are using WeightWatcher and trying to analyze models trained with Muon, feel free to reach out for help. #talkToChuck 𝗔𝗱𝗮𝗺𝗪 𝘃𝘀. 𝗠𝘂𝗼𝗻: α, 𝗢𝘃𝗲𝗿𝗳𝗶𝘁𝘁𝗶𝗻𝗴, 𝗮𝗻𝗱 𝗠𝗲𝗺𝗼𝗿𝗶𝘇𝗮𝘁𝗶𝗼𝗻. A quick update on our WeightWatcher experiments comparing AdamW with the newer Muon optimizer. We trained a single-head nanoGPT model with both optimizers across five seeds. We then used WeightWatcher to study the weight-matrix spectra and several forms of memorization. The top plot shows the average raw α. AdamW produces relatively stable α values. Muon is much noisier. Its ESDs vary substantially across layers, which makes α harder to estimate reliably. But the average hides something important: 👉 Individual layers do fall below α = 2. Our spectral/RG theory predicts that this regime can be associated with pathological, example-specific learning. So we tested that prediction directly. We inserted random sequences into the training set and measured whether the model remembered them using exact extraction, token recall, teacher-forced recall, and canary exposure. The result was striking. • AdamW memorized the planted sequences strongly. Earlier in training it reached ~96% exact recall of 32-token canaries. • Muon strongly suppressed this kind of memorization. At 10,000 steps, teacher-forced recall was ~62% for AdamW versus <1% for Muon. There is also a surprise: some Muon layers have α < 2 as well. So α < 2 by itself is not sufficient to explain memorization. This is why we recommend examining individual WeightWatcher layers and their ESDs—not simply average α. Muon appears to suppress memorization while also producing unusual ESDs that make conventional WeightWatcher analysis more difficult. We are now running Muon much longer to understand how these spectra evolve. If you are using WeightWatcher and trying to analyze models trained with Muon, feel free to reach out for help. #talkToChuck
4
16
133
45,386
𝗖𝗿𝗲𝗮𝘁𝗶𝗻𝗴 𝗮 𝗟𝗶𝗾𝘂𝗶𝗱 𝗣𝗼𝗰𝗸𝗲𝘁 𝗜𝗻𝘀𝗶𝗱𝗲 𝗮 𝗖𝗿𝘆𝘀𝘁𝗮𝗹: 𝗠𝘆 𝗘𝗮𝗿𝗹𝘆 𝗪𝗼𝗿𝗸 𝗼𝗻 𝗠𝗼𝗻𝘁𝗲 𝗖𝗮𝗿𝗹𝗼 🔬 A recent conversation about Monte Carlo simulations reminded me of my undergraduate thesis work with Sherwin J. Singer. I developed a new Monte Carlo technique to test whether melting could begin around a vacancy in a crystal. Monte Carlo is one of the most widely used methods in computational science. Its modern development at Los Alamos grew out of neutron-transport problems: sampling how neutrons scatter, diffuse, become absorbed, and produce further neutrons through fission. That same research community brought Monte Carlo into statistical mechanics. In 1953, Metropolis and colleagues introduced a method for sampling molecular configurations according to their Boltzmann probabilities. Their first calculations studied the equation of state of hard disks. The method became a foundation for simulations of fluids, solids, and phase transitions. My undergraduate work used that framework to investigate a specific question: could a vacancy support a small liquidlike pocket before the surrounding crystal melted? An ordinary simulation could miss such a pocket. A free-energy barrier might prevent it from forming within the available simulation time. Failure to observe the pocket would then leave its stability unresolved. 𝗜 𝗰𝗼𝗻𝘀𝘁𝗿𝘂𝗰𝘁𝗲𝗱 𝘁𝗵𝗲 𝗽𝗿𝗼𝗽𝗼𝘀𝗲𝗱 𝘀𝘁𝗮𝘁𝗲 𝗱𝗶𝗿𝗲𝗰𝘁𝗹𝘆. We divided the system into two regions: • Outer atoms could move within their assigned Wigner–Seitz cells. • A central region containing vacancies remained free to rearrange. We began at low density and compressed the system. The constraints caused the surrounding atoms to order first, creating an artificially stabilized crystal around a liquidlike pocket. We then tested whether that pocket could remain disordered. 𝗜𝘁 𝗿𝗲𝗳𝗿𝗼𝘇𝗲. The pocket reordered even before we reached the bulk freezing density. In our inverse-sixth-power model, vacancies and clusters of up to four vacancies did not sustain premelting. Preparing the liquidlike state directly addressed the concern that a formation barrier had hidden it from ordinary simulations. The methodological idea still interests me: use constraints to construct configurations that the unconstrained system would not ordinarily sustain, then test their stability through Monte Carlo sampling. The work appeared in our 1991 paper, “Behavior of point defects in a model crystal near melting,” in Physical Review B. Read the paper
9
501