assistant professor of computer science @hseas, learning theorist, 🎹

Pinned Tweet
🚨 We finally have a better understanding of how the different paradigms for discrete diffusion compare! In new work with Liye Wang, we prove a quadratic separation between the parallel generation capabilities of uniform and Gaussian diffusion relative to unmasking! 1/
4
25
177
13,844
Finally here - Submitted PHA-4 to arXiv: drive.google.com/file/d/1m3g… Sampling SK all the way to beta < 1 building on the free-probability and cavity interpolation theory developed in PHA-3, with the amazing @oldheneel and @jtnshi 1/n
1
5
15
1,266
Sitan Chen retweeted
There is growing interest in generating discrete data via diffusions. If you have an approach, consider benchmarking it on generating uniform solutions of (random) k-SAT. It is a clean example in which continuous beat discrete diffusions and the diffusion process matters.
3
19
197
12,002
🚨 We finally have a better understanding of how the different paradigms for discrete diffusion compare! In new work with Liye Wang, we prove a quadratic separation between the parallel generation capabilities of uniform and Gaussian diffusion relative to unmasking! 1/
4
25
177
13,844
Finally, please also check out these amazing concurrent works by @dmitrievdaniil7, Zhihan Huang, and Yuting Wei, as well as by Martin Wainwright! arxiv.org/pdf/2608.23554 arxiv.org/pdf/2608.28949 Lots of important and surprisingly crisp mathematical questions in this space! 7/7
1
1
12
695
Sitan Chen retweeted
If you think no one notices: they do. It's really, really obvious.
every day I feel like shouting into the void: if you wish to be known as a professional of high taste, do not post AI slop the game is long, and you risk your reputation sinking rapidly amongst the people who matter. Thank you for your attention to this matter!
3
108
19,289
📢Aug 31 (Mon): From Interface to Inference: Eliciting Any-Order Inference from Any-Order Models 🤔Many discrete reasoning tasks, such as code generation, are inherently non-causal. Programmers naturally move between high-level structure and local details in a process called any-order inference. While masked diffusion models offer a native any-order prediction interface, maximizing their potential is difficult due to a core limitation: 1️⃣ This any-order interface does not automatically yield actual any-order inference 2️⃣ Fixed-canvas, token-level models suffer from "positional uncertainty"—they may know what semantic component should appear, but not where to place it 🔑To address the interface-inference gap, the authors propose two complementary approaches designed to natively support any-order inference without relying on hand-designed mechanisms. 🔧The framework tackles positional uncertainty through two methods: (1) Insertion-based masked diffusion (building on FlexMDM), which relaxes fixed-position commitments via insertions to enable generation across non-contiguous regions, and (2) Latent-space masked diffusion (LatentMDM), which shifts prediction to coarser semantic segments to enable search over latent generation orders. 📈The resulting models—a 7B FlexMDM trained for Python coding and a 125M LatentMDM for GSM8K—successfully induce distinct any-order inference behaviors. The study demonstrates that overcoming positional uncertainty through these approaches significantly improves downstream performance. This Monday, the joint first authors Seunggeun Kim (@seunggeun_kr), Jaeyeon Kim (@Jaeyeon_Kim_0), Taekyun Lee (taekyunl.github.io), and Yuyuan Chen (yuyuanchen0.github.io) will present their recent paper on the FlexMDM and LatentMDM frameworks.
2
10
19
7,625
Sitan Chen retweeted
With the start of the semester just a week away, here’s what I decided to say to my upper division math course about how AI will or won’t change what they’re learning… The Enduring Value of Math, in an Age of AI francissu.substack.com/p/the…
4
53
212
76,567
Just because a model can decode tokens in any order doesn’t mean it can truly reason in any order.. Check out Jaeyeon’s awesome thread about our latest work that resolves this conundrum and finally realizes the promise of any-order models!
Excited to share what we have worked on for the past 6 months. Can we really call masked diffusion an "any-order" model? In our new work, we argue that off-the-shelf masked diffusion models have any-order predictive power but fail to reason in a 'genuinely any-order fashion'. We introduce two models that do: LatentMDM and 7B-scale FlexMDM, extending our prior work. LatentMDM, our new latent-space masked diffusion, outperforms AR (with KV caching) on TinyGSM, at a matched inference budget. To the best of our knowledge, LatentMDM is the first model on the same setup to do so. Joint work with co-first authors @seunggeun_kr, @tklll_tx, and Yuyuan Chen, and advisors @du_yilun, @ShamKakade6, and @sitanch.
2
20
3,075
Sitan Chen retweeted
Rapidly advancing AI models have sparked breakthroughs in long-standing math problems. From solving problems at the International Math Olympiad to disproving decades-old conjectures - some in the math world are now questioning whether they have a future in the field. But not Professor Shayan Oveis Gharan, who won the 2026 Abacus Medal for his work on the theory of algorithms. He believes the power of AI is harnessed not as replacement for human mathematicians, but as a partner. CBS’s Kamal Afzali explains.
10
21
76
57,760
Sitan Chen retweeted
Very honestly, the current status of research in this area is more sophisticated than this result. This way of communicating science can make serious and long lasting damage. [This is not to say that the result isn’t new or that AI models aren’t amazing.]
👀 GPT-5.6 Sol and Fable 5 just cracked a 25 year old open problem in wireless communication theory
12
35
345
53,493
Sitan Chen retweeted
1/ Excited about our new paper w/ @HongYeHu1: Provably Efficient Self-Calibrating Quantum Fault Tolerance. The syndrome bits of QEC code are secretly a calibration signal — and we prove that one can ride them to keep the hardware in tune without ever pausing the computation. ⚛️
2
2
22
2,524
Excited to be a part of this amazing community!
We are thrilled to announce the appointments of two new associate faculty members: Nada Amin and Sitan Chen — pioneering scientists who joined the #KempnerInstitute on July 1! Read more: bit.ly/4vmnQHZ @sitanch @hseas #AI #ML
2
24
4,065
A poster session you won't want to miss this afternoon! And stay tuned for some exciting new dLLM results that we have in the works :)
1
2
11
2,576
For those lucky enough to be in Seoul right now, make sure to check out Adil's oral presentation in a few hours on our work, joint with Khashayar Gatmiry, on high-accuracy diffusion sampling!
Today at #ICML 2026, I'll be presenting our paper on score-based sampling that achieves both high-accuracy and dimension-free guarantees. ORAL Wed, Jul 8, 2026 • 10:30–10:45 AM KST • Hall C POSTER Wed, Jul 8, 2026 • 2:30–4:15 PM KST • Hall A #2507 Please join us if you're around!
3
37
16,986