Ass. prof. of Machine Learning. PI of Generative Memory Lab (@DondersInst). Generative diffusion and statistical physics. AI realist.

Nijmegen, Nederland
1/?) As promised to Sander Dieleman (@sedielem), we’re finally excited to share: Towards Closing the Autoregressive Gap in Language Modeling via Entropy-Gated Continuous Bitstream Diffusion We show that continuous diffusion can achieve very strong language modeling performance when operating directly on bitstreams, outperforming masked and uniform diffusion baselines, and essentially matching autoregressive models under our evaluation settings.
6
37
237
33,206
Luca Ambrogioni retweeted
We will be presenting this at NeurIPS in Paris in December! 🇫🇷 📉
🚨 New paper: Introducing MIND (Monge Inception Distance) Everyone agrees that FID is broken, requires too many samples, slowing down evals. MIND requires 10x fewer samples, is more robust, faster to compute. Our new drop-in replacement for evaluating generative models. 🧵👇
2
7
80
5,091
Luca Ambrogioni retweeted
This is going to completely revolutionize fields like QCD phenomenology.
New on the Science Blog: Yes, Claude can do Nine Loops. Theoretical physicists predict how particles behave using formulas called scattering amplitudes. These are notoriously hard to compute, so researchers work with layers of increasingly fine corrections called “loops”—each added loop makes the answer more precise but takes exponentially more computation. Most calculations stop at two or three loops. Eight loops was the previous record in a simplified model physicists use as a testing ground (planar N=4 super-Yang-Mills), set by SLAC's Lance Dixon and collaborators. Last month, physicist and science writer @4gravitons issued a challenge: could an AI push past eight loops in this model, using only the compute budget an academic could reasonably access? Given a single prompt describing the nine-loop problem, Claude ran largely unsupervised for days in Claude Science and solved it using methods developed by Dixon and his colleagues, at a total cost of a few thousand dollars. Dixon independently verified the result, and von Hippel wrote about the experience for our blog. Read more: anthropic.com/research/yes-c…
5
7
134
10,199
Luca Ambrogioni retweeted
History of diffusion model foundations A slide from a recent tutorial I gave. Genuinely curious if I missed anything
5
3
50
3,150
Luca Ambrogioni retweeted
Happy to share that our paper was accepted to #NeurIPS2026 as a spotlight! 🎉
A while ago we presented an early version of our work on spherical flows for categorical data at @diffusion_llms. The paper has grown a lot since, so it felt like time for a thread...
2
10
24
1,767
Wow...
A very very bad decision.
1
7
961
Luca Ambrogioni retweeted
I think there are many questions related to AI on which there are no real experts. That is often a better starting point for discussion than referring to what experts think, where the relevant expertise is on much narrower and more technical questions.
16
10
81
10,134
Luca Ambrogioni retweeted
It is important to center the fact that Claude built on the approach that my colleague Lance developed, used tools that the field has developed, and (from what I have been told) used results in some of our recent papers. It’s a huge accomplishment, but should be framed as a combination of human and AI contributions. (To @AnthropicAI’s credit, I think the article does a pretty good job of this.) The fact that Claude can take a single people and figure out how to integrate the knowledge and use the tools as amazing, but it’s not coming out of the vacuum. This is an important opportunity for us to start to frame these developments, recognizing the human contributions and the power of these new tools. This is not an exemplar of the bitter lesson!
holy shit. claude just pushed a theoretical physics calculation beyond the previous record after working on it largely by itself for days. >previous record: 8 loops >claude reached 9 >largely unsupervised >wrote its own code >found two different ways to calculate it >both gave the same result >107,053 nonzero coefficients matched physicists then spent weeks checking the work this is the shit i’ve been waiting for. science is about to get fucking weird.
6
31
203
18,150
Luca Ambrogioni retweeted
I am writing a super accessible explainer of AI (for non-CS folks). Anyone wanna test drive?
7
6
43
4,159
Luca Ambrogioni retweeted
⭐️Continuous Diffusion Scales Competitively with Discrete Diffusion for Language Has been accepted at NeurIPS 2026! See you at the conference!
📢Excited to share our new paper: Continuous Diffusion Scales Competitively with Discrete Diffusion for Language We introduce RePlaid 🌊, a continuous diffusion language model (DLM) with 🏅Discrete likelihood bound 🏅Scaling laws competitive with SOTA discrete DLMs How? Dive in👇[🧵1/12] Paper: arxiv.org/abs/2605.18530 Work done with my amazing collaborators: @WeiGuo01 @ShuibaiZ69721 @ssahoo_ @YongxinChen1 @ArashVahdat @MardaniMorteza @jwthickstun
3
13
81
4,869
Luca Ambrogioni retweeted
Jürgen Schmidhuber is joining Sakana AI as Chief Scientific Advisor. @SchmidhuberAI pioneered meta-learning, recursive self-improvement, and world models back in the 1990s, when compute was a million times more expensive. He has been thinking about machines that improve themselves since before compute was cheap enough to make it practical. These ideas inspired the Darwin Gödel Machine and The AI Scientist. Our RSI Lab in Tokyo, now under Jürgen’s guidance, is working on agent-native world models and recursive self-improvement for physical AI.
Sakana AI welcomes Jürgen Schmidhuber as Chief Scientific Advisor. sakana.ai/schmidhuber/ Sakana AI is incredibly proud to announce that Jürgen Schmidhuber, universally recognized as the father of modern AI, is officially joining Sakana AI as Chief Scientific Advisor. For nearly four decades, Jürgen has explored how machines can learn to learn. His foundational work in the 1990s drove core advancements in deep learning and established early frameworks for world models. Crucially, his pioneering innovations in meta-learning opened the very path toward recursive self-improvement. These ideas have already shaped our own research, from the Darwin Gödel Machine to The AI Scientist. Now Jürgen will help guide our newly formed RSI Lab, whose objective is to trigger a compounding cycle of scientific discovery aimed at improving machine intelligence. We are assembling a critical mass of world-class experts in Tokyo to make this a reality. Welcome, @SchmidhuberAI !
70
76
965
68,685
Luca Ambrogioni retweeted
NeurIPS decisions are out
7
3
65
19,300
Luca Ambrogioni retweeted
A reminder that style guides have become a vehicle for the advancement of political agendas.
Artificial intelligence systems do not think, feel, want or understand. Avoid language that gives them human characteristics. This is called anthropomorphizing, when we ascribe human traits, emotions or behaviors to non-human things, such as animals or inanimate objects. Instead, explain what a system does, how well it performs, who built it and who could be affected by it. apnews.com/article/openai-sa…
3
1
38
3,035
Have an AI ever asked you a question where you felt it was driven by its own curiosity?
2
1
1
694
Luca Ambrogioni retweeted
There is growing interest in generating discrete data via diffusions. If you have an approach, consider benchmarking it on generating uniform solutions of (random) k-SAT. It is a clean example in which continuous beat discrete diffusions and the diffusion process matters.
3
19
198
11,990
I am both terrified and extremely excited at the idea of AI developing true intentionality. It is something I really want to experience, while I part of me hopes that it will never come. So far, I honestly do not know if it is even possible with our current technology or if it is right around the corner. Exciting times.
5
2
559
To make it clear, in the context of this spicy take I am considering complex analysis, matrix algebras, lie group theory and differential geometry as a little more advanced than calculus I hope you are not triggered by my definition :p In the scope of modern math those are elementary topics (as far as the physically useful part is concerned)
People wonder about the unreasonable effectiveness of math, but the real question is about the unreasonable ineffectiveness of any math that is more than just a little more advanced than Calculus
10
34
5,186
Luca Ambrogioni retweeted
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
2,223
5,277
53,327
9,755,676
Luca Ambrogioni retweeted
Paper thread! I recently read this paper from some FAIR colleagues and NYU folks in depth, so making a thread while it's fresh in my head. The main point is seeing how far on the SMALL side we can go in scaling laws without losing fit or predictability, and what it takes?
12
58
704
43,608
People wonder about the unreasonable effectiveness of math, but the real question is about the unreasonable ineffectiveness of any math that is more than just a little more advanced than Calculus
24
9
159
17,895
Luca Ambrogioni retweeted
An amazing work explaining the difference between continuous and discrete diffusion in terms of the size of their critical decision windows, reflecting their symmetry breaking dynamics! This deserves another Youtube thumbnail!
Replying to @sitanch
We instead show that the real reason is that critical windows for masked diffusion are quadratically narrower than critical windows for uniform/Gaussian! 3/
2
3
24
2,293