PhD student at @MIT_CSAIL working on self-improvement and synthetic data || prev at @GoogleDeepMind @AIatMeta

Can a model learn to break its own reasoning plateau? In our new paper, we show that LLMs can be taught with meta-RL to generate their own "stepping stones" that kickstart learning on hard math problems (0/128 success rate) where direct RL fails. Paper 📝: arxiv.org/abs/2601.18778 Blog post 🌐: ssundaram21.github.io/soar/ (1/n)
20
114
696
125,985
Awesome work as always by @SophieLWang w/ @_AmilDravid, with very neat and intuitive insights about how mid-training data affects downstream reasoning!
Can “chicken” make a base model reason? 🐔 Yes! Why? With the right first tokens, base models can match RL in reasoning. We trace this to learned training data associations and use a simple data edit to make “chicken” elicit reasoning too. blog: sophielwang.com/cues
1
13
687
Synthetic data stays winning 🚀 now applied to xrays! Love the demos and figures too
X-rays are everywhere in medicine, but extracting general, quantitative anatomical information from them remains remarkably difficult. Announcing FleXray: an open-source, open-weight model for zero-shot anatomical segmentation across the human body—from different regions, viewing angles, and acquisition settings. FleXray learns from large-scale synthetic supervision generated from 3D anatomy, then transfers directly to real clinical X-rays. Try it now on your own X-rays in our in-browser demo! 🧵 1/N — project page, demo, weights, code & paper below ↓
1
7
925
Shobhita Sundaram retweeted
More efficient test time LLM reasoning with learned concepts: Check out our new paper "Beyond Repeated Sampling"! arxiv.org/abs/2609.26704
Repeated sampling is the default way to scale LLM reasoning at test time. But token level noise often produces many near duplicate attempts that follow the same high level idea. 🧵 How can we cover more of the solution space without sacrificing throughput?
2
18
58
4,972
Shobhita Sundaram retweeted
Can we tell whether data domains cooperate or compete during pretraining? Adding code to the mix makes models better at math, while some other combinations hurt each other. We call this data synergy. Turns out you can incorporate data synergy into scaling laws and estimate it 🧵
4
33
208
23,670
I'll be at #ICML2026 to present our *spotlight* work! TLDR: LLMs can learn to self-generate curricula for problems they can't yet solve, using self-play with meta-RL. Please reach out to chat about self-improving agents, synthetic data & environments, curriculum learning, or anything else! We've updated the paper with some fun additions ⬇️
Can a model learn to break its own reasoning plateau? In our new paper, we show that LLMs can be taught with meta-RL to generate their own "stepping stones" that kickstart learning on hard math problems (0/128 success rate) where direct RL fails. Paper 📝: arxiv.org/abs/2601.18778 Blog post 🌐: ssundaram21.github.io/soar/ (1/n)
7
21
150
19,142
3. Our main result extends to bigger models (Llama-3.1-8B-Instruct).
1
1
3
246
Blog post and paper: ssundaram21.github.io/soar/ 📍Presenting our poster on Weds, Jul 8, 5PM KST!
1
4
202
Shobhita Sundaram retweeted
Super excited to be able to release this project we've been working on. The largest open robotics dataset to date and policies capable of dexterous manipulation!
Introducing ABC: open data, training, and infrastructure for robotics. We release the largest teleop dataset to date, and extensively investigate design decisions, pretraining, and post-training techniques. @arthurallshire @Cinnabar233 @adamrasb @redstone_hong @davidrmcall
1
9
54
9,969
Super cool work from @juliachae_ that fills an important gap in image similarity metrics!
Excited to share ID-Sim, our identity-focused similarity metric, presenting at #CVPR2026 this week in Denver! 🎉 Humans are remarkably good at distinguishing highly similar objects across different contexts. We asked: can we train a metric that does the same?
3
6
1,413
Shobhita Sundaram retweeted
Our team at @AIatMeta is excited to announce ATLAS: one of the largest automated formalization efforts to date. ATLAS contains Lean 4 formalizations of both statements and proofs from 25+ mathematics textbooks, spanning dozens of domains, for a total of 500k lines of code. We are also releasing a flexible formalization harness and a companion paper. External contributions are welcome! Joint work spearheaded by our amazing PhD student Ahmad Rammal (@Ahmad3Rammal), together with Niket Patel (@niketnpatel ), Fabian Gloeckle (@FabianGloeckle), Amaury Hayat (@Amaury_Hayat), Remi Munos (@MunosRemi), Julia Kempe (@KempeLab), Vivien Cabannes, and myself from @AIatMeta, @NYUDataScience , and Ecole des Ponts. This is an ongoing effort; more details in the thread below. (1/9)
29
87
440
377,549
Shobhita Sundaram retweeted
What is the right data mix, and how do we find it as the data keeps changing? This is a core, unsolved problem in continual learning. To tackle it, we built a data mixing algo that works everywhere — pretraining, midtraining, instruction tuning Introducing: On-Policy Mix 🧵1/6
6
57
330
53,126
Shobhita Sundaram retweeted
"The Truth Lies Somewhere in the Middle (of the Generated Tokens)" In autoregressive language models, mean pooling hidden states across generation yields better representations than any token alone. project page: sophielwang.com/tokens w/ @phillip_isola and @thisismyhat
11
71
486
64,174
Shobhita Sundaram retweeted
As a geometric ML researcher, I noticed pseudoscalars don’t get enough attention! Read on to see what pseudoscalars can do for you.
1
2
3
460
LLMs can learn to self-generate curricula for hard problems that they can't yet solve! Using meta-RL, with rewards grounded in learning progress, models produce their own stepping stones that kickstart learning on hard problems where direct RL plateaus. Poster at the ICLR RSI workshop today!
Can a model learn to break its own reasoning plateau? In our new paper, we show that LLMs can be taught with meta-RL to generate their own "stepping stones" that kickstart learning on hard math problems (0/128 success rate) where direct RL fails. Paper 📝: arxiv.org/abs/2601.18778 Blog post 🌐: ssundaram21.github.io/soar/ (1/n)
1
24
168
16,873
Shobhita Sundaram retweeted
New blogpost on tokenizing non-sequential data! Language has sequential structure, which gave rise to the next-token prediction paradigm of LLMs. But we increasingly use LLMs for data without inherent order (e.g. images, molecules, sets). What does “next token” mean here? (1/7)
6
28
271
27,580
Shobhita Sundaram retweeted
Introducing GASP😮: Guided Asymmetric Self-Play for Coding LLMs We address the goal-agnostic behavior of current asymmetric self-play methods. Key idea: guide the teacher with hard real-data goalposts; first an easier lemma, then a harder lift from the lemma as stepping stones 🧵
1
11
71
11,934
Shobhita Sundaram retweeted
Simply adding Gaussian noise to LLMs (one step—no iterations, no learning rate, no gradients) and ensembling them can achieve performance comparable to or even better than standard GRPO/PPO on math reasoning, coding, writing, and chemistry tasks. We call this algorithm RandOpt. To verify that this is not limited to specific models, we tested it on Qwen, Llama, OLMo3, and VLMs. What's behind this? We find that in the Gaussian search neighborhood around pretrained LLMs, diverse task experts are densely distributed — a regime we term Neural Thickets. Paper: arxiv.org/pdf/2603.12228 Code: github.com/sunrainyg/RandOpt Website: thickets.mit.edu
91
463
3,171
790,194
Shobhita Sundaram retweeted
Can language models learn useful priors without ever seeing language? We pre-pre-train transformers on neural cellular automata — fully synthetic, zero language. This improves language modeling by up to 6%, speeds up convergence by 40%, and strengthens downstream reasoning. Surprisingly, it even beats pre-pre-training on natural text! Blog: hanseungwook.github.io/blog/… (1/n)
46
257
1,662
260,553