Assoc Prof at UC San Diego @ucsd_cse, AI researcher

San Diego, CA
Taylor Berg-Kirkpatrick retweeted
Stanford Prof. @chrmanning says universities are losing AI professors to frontier labs because they still fund computer science like it’s 2005, when a laptop was enough: "Universities have just been too flat-footed in dealing with the way the computing and AI world has changed. There are really two problems." "Over the last couple of decades, universities have become enormously more bureaucratic and administrative, and that makes them a less good research environment than they used to be in the 20th century." "Universities got used to this nice period around 1990 to 2010, when computer scientists needed barely any money for their research. You bought a laptop or a PC, and it was fine. The reality now is that computer scientists, particularly in AI, need a lot of money." "It's established that if you're hiring a new faculty in physics or chemistry, you need some millions of dollars, but it hasn't trickled down that people in AI need that too. There are too many professors leaving academia." "Stanford made a start with the Marlowe cluster of 256 H100s. That's an okay baby step, but there probably needs to be about 10 times that much." @stanfordnlp @Stanford
11
50
308
99,551
Taylor Berg-Kirkpatrick retweeted
This is exactly what I hear again and again when talking to experts. Biorisk is fake. The problem is not designing a dangerous pathogen -- AI is good at that -- it is creating it physically to make it dangerous. And that is what AI cannot do and is easy to safeguard against
I must be among an extremely small group of people (n=1?) that have both 1) trained a frontier LLM and 2) designed and synthesized custom viruses in a lab with my own two hands. And I think that the takes on AI killing us all by creating dangerous viruses is total bogus.
28
19
216
21,892
Taylor Berg-Kirkpatrick retweeted
RL is expensive, so every step should count. ~1 yr ago, we showed xent SFT isn’t the best way to prepare for RL (arxiv.org/abs/2510.15020). Now, we propose TailSFT (arxiv.org/abs/2608.25756), a lightweight + principled way to directly improve coverage and get better post-RL perf.
8
82
685
155,057
Taylor Berg-Kirkpatrick retweeted
Prior to 2012 just about nobody in the core software industry (MSR doesn’t count) paid much attention to academic CS publishing. Also, prior to 2012, academics actually read papers and judged you more on the quality than the quantity of those papers. Despite the low cost of producing shit papers, even then, the system kinda worked mainly because there was little incentive to spray slop. The groundswell of interest in the outputs of CS academia broke the system. The reward for “has a NeurIPS”, divorced from the “I read this paper and want to interact with this mind” broke the system. Perhaps all that it takes for the system to rebuild is for it to burn to the ground. Once everyone with money and power thoroughly decides to ignore papers, the incentives to spam will evaporate. From the ashes of the scorched landscape of academic publishing a few fresh shoots will sprout. These will be weirdos with ideas that care about their ideas and don’t care if there’s a pot of gold at the end of the review cycle. Open science and academic inquiry will bloom again, until overgrowth and an untimely drought ignite the next wildfire.
The cost of producing low quality papers has gone to almost zero, so to maintain the publication system we need to increase the cost of submitting low quality papers. The easiest way to do this is to impose reputational cost for producing low quality work. How to do this?
7
14
178
20,067
Taylor Berg-Kirkpatrick retweeted
this claim, that AI has a real risk of killing all of humanity in the near* future, has been stated regularly without evidence for over a decade. it hasn't even come close in that time. is it reasonable? two sanity check comparisons: climate change, and life w/o internet. 1/n
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
26
18
197
22,665
Taylor Berg-Kirkpatrick retweeted
🤔4B params are not enough for a good coding agent... Think again! We show how to get a 4B model to be a good agentic coder without distillation! Introducing FrogNano 🐸 Check the thread for more information 👇
4B parameters. ZERO distillation. 61.5% on SWE-bench Verified 🤯 Meet FrogNano 🐸: Qwen3.5-4B post-trained purely with RL on synthetic tasks from TaskPilot. Just 5 iterations × 300 tasks. Who said coding agents have to be huge? 🐸
6
8
172
12,008
Taylor Berg-Kirkpatrick retweeted
This is cool, but a bit confused how Table 1 can claim that our work, Parcae, doesn’t compare isoFLOPs results between looped and fixed depth? A major contribution (section 5.2) of our paper is that looping is compute optimal under isoFLOP and isoParams. arxiv.org/abs/2604.12946
3
1
16
1,031
Taylor Berg-Kirkpatrick retweeted
It's exciting to see the rumors that Astra is a looped model. For those interested in the math on how to stably train them, I would check out our work Parcae. We also show that looping follows predictable scaling laws! arxiv.org/pdf/2604.12946
new: OpenAI & others quietly using loop transformers that don't show their 'thinking' when scaled up a leap forward on performance, but sparking concerns inside & outside OpenAI re: security as this takes off
3
11
53
10,443
Rumor has it OpenAI's Astra is a looped model. If you want the math on how to actually train one of these without residual explosion and loss spikes, we (@hayden_prairie @zacknovack @realDanFu) worked it out in Parcae (spectral norm constraints on the injection params, per-sequence loop sampling, scaling laws). Curious whether OpenAI hit the same instabilities and fixed them the same way. arxiv.org/abs/2604.12946
new: OpenAI & others quietly using loop transformers that don't show their 'thinking' when scaled up a leap forward on performance, but sparking concerns inside & outside OpenAI re: security as this takes off
1
1
22
21,809
Taylor Berg-Kirkpatrick retweeted
In a new episode of stealing parts of production LLMs, we found a way to steal hidden model architecture and expose undocumented inference optimizations through ordinary streaming APIs. On Gemini Flash 2.5, latency jumped 3.2× at 130k tokens, revealing speculative decoding and a hidden 128K draft-model context window. sadegh-majidi.github.io/leak… 🧵
20
26
209
13,749
Taylor Berg-Kirkpatrick retweeted
New blog post: continuous diffusion for language is back! This research direction receded into the background for a while, but as of this year, it is once again a hot topic. I wrote down a historical perspective and some thoughts on the recent revival. sander.ai/2026/08/24/continu…
22
135
767
91,988
Taylor Berg-Kirkpatrick retweeted
Rigid, hand-designed sampling schedules become inaccurate during autoregressive video generation as errors accumulate. We introduce Equilibrium Forcing (EqF), a modular framework for video generative modeling that removes noise level conditioning to enable adaptive inference. 🌀
5
28
133
28,342
Taylor Berg-Kirkpatrick retweeted
I'm so happy to be able to announce the first general purpose sign-language-to-text translation (SL2T) model from my team that's out today and powering a new ASL input feature on @Android phones (and don't worry, we're not stopping at ASL).
20
68
362
110,288
Taylor Berg-Kirkpatrick retweeted
Replying to @huskydogewoof
@huskydogewoof is truly one of the looped model GOATs, awesome work here! AND they used Parcae-style stable input injection too🤘
🔥 𝐍𝐞𝐰 𝐛𝐥𝐨𝐠: 𝐓𝐨𝐰𝐚𝐫𝐝𝐬 𝐋𝐨𝐨𝐩𝐞𝐝 𝐌𝐨𝐝𝐞𝐥𝐬 𝐃𝐨𝐧𝐞 𝐑𝐢𝐠𝐡𝐭 — 𝐏𝐚𝐫𝐭 𝐈 Looped models reuse the same weights across depth, promising a better compute–parameter trade-off, especially for reasoning. 𝐁𝐮𝐭 𝟏) 𝐝𝐨 𝐭𝐡𝐞 𝐠𝐚𝐢𝐧𝐬 𝐬𝐮𝐫𝐯𝐢𝐯𝐞 𝐰𝐡𝐞𝐧 𝐛𝐨𝐭𝐡 𝐭𝐫𝐚𝐢𝐧𝐢𝐧𝐠 𝐚𝐧𝐝 𝐢𝐧𝐟𝐞𝐫𝐞𝐧𝐜𝐞 𝐅𝐋𝐎𝐏𝐬 𝐚𝐫𝐞 𝐦𝐚𝐭𝐜𝐡𝐞𝐝? 𝟐) 𝐀𝐧𝐝 𝐰𝐡𝐢𝐜𝐡 𝐚𝐫𝐜𝐡𝐢𝐭𝐞𝐜𝐭𝐮𝐫𝐚𝐥 𝐜𝐡𝐨𝐢𝐜𝐞𝐬 𝐚𝐜𝐭𝐮𝐚𝐥𝐥𝐲 𝐦𝐚𝐭𝐭𝐞𝐫? We run 𝐚𝐩𝐩𝐥𝐞𝐬-𝐭𝐨-𝐚𝐩𝐩𝐥𝐞𝐬 ablations spanning Ouro to Huginn. Huginn performs better overall, with the largest gains coming from the loop-in-the-middle (sandwich) design and input injection, though they provide different benefits. Trained on 𝟓𝟎𝟎𝐁 tokens, an 𝟖𝐁-𝐀𝟎.𝟖𝐁 Huginn MoE approaches or surpasses a 𝟑𝟐𝐁-𝐀𝟑.𝟐𝐁 feedforward MoE on several reasoning benchmarks, including GSM8K (83.6% vs. 80.8%), while using 𝟕𝟓% 𝐟𝐞𝐰𝐞𝐫 resident parameters under 𝐦𝐚𝐭𝐜𝐡𝐞𝐝 𝐭𝐫𝐚𝐢𝐧𝐢𝐧𝐠 𝐚𝐧𝐝 𝐢𝐧𝐟𝐞𝐫𝐞𝐧𝐜𝐞 FLOPs. More details and the blog link in the thread ↓
1
1
9
859
Taylor Berg-Kirkpatrick retweeted
why can the china labs build glm-5.2, kimi k3, and many more to come? it is because of the openness. not just the open weights but the whole ecosystem. most of the work done in the china labs is carried by interns. i met brilliant undergrad and graduate interns who deeply understand the model training details, and they are 100x more open to share. that means the talent that knows how to train llms in china is 100x greater in number than the talent in the us, and it is growing in contrast, the us ai ecosystem is too closed. frontier labs do not hire interns. i know brilliant phd students at stanford, berkeley, and so on. they struggle to get an internship and the compute to train a properly sized model. most of the secret recipes are locked away by a very small group of privileged researchers it is not about china or the us. it is about open and closed science. the fact is that every average cs student can learn how to train an llm. they just need the opportunity. labs should be more open and hire more interns, like how deepmind and fair did in the pre-llm era
33
117
946
263,782
🎉 Thrilled to share we had 4 papers accepted to COLM 2026! Huge congrats to Hayden Prairie @hayden_prairie, Ivan Lee @ivn1e (lead on two!), Yasaman Jafari @yasjafarii , Zixian Wang, Cheng Yang @ChengYANG_yc , and Zachary Novack @zacknovack — 🧵 below with all the papers: 1/ "Parcae: Scaling Laws For Stable Looped Language Models" Hayden Prairie, w/ Zachary Novack, in collab with Dan Fu @realDanFu Viewing layer looping as a dynamical system yields stable looped LMs, and a new predictable scaling axis at constant memory. 📄 arxiv.org/abs/2604.12946 2/ "Studying the Soupability of Documents in State Space Models" Yasaman Jafari, w/ Zixian Wang & Leon Bergen Encode docs separately with Mamba2, then average the SSM states into one "soup": better multi-doc QA at a fraction of the inference cost, scaling to 256 docs. 📄 arxiv.org/abs/2505.24033 3/ "The Format Tax: Measuring and Mitigating the Cost of Structured Output" Ivan Lee, w/ Loris D'Antoni Forcing JSON/XML output hurts reasoning. The culprit is prompt-side distribution shift, not constrained decoding. Reason freely, then reformat: most of the accuracy comes back. 📄 arxiv.org/abs/2604.03616 4/ "Optical Context Compression Is Just (Bad) Autoencoding" Ivan Lee, w/ Cheng Yang Vision tokens aren't magic: simple baselines like mean pooling match DeepSeek-OCR-style compression on reconstruction and beat it for language modeling. 📄 arxiv.org/abs/2512.03643 #COLM2026
2
11
70
7,320
Taylor Berg-Kirkpatrick retweeted
Lol ok this auditorium is huge haha No pressure, go humans!! 💪
If you are around, come to the Auditorium (3rd floor, on the north side of COEX) at 3:15pm and watch the (anti)debate between @lschmidt3 and I on: wether we should keep paying human researchers post superhuman AI!
4
160
15,858
Taylor Berg-Kirkpatrick retweeted
#3105!! Join us!
🚨New paper to level up your 🦞#Clawdbot ?! Bots are now posting your sensitive info in real time. But privacy research is a desert with no data to train better models. That's about to change Enter 🏝️Privasis, the oasis where you can train strong privacy-forward AI with scale✨
3
5
169
15,817
Taylor Berg-Kirkpatrick retweeted
If you're interested in unified multimodal models, come say hi and check out our poster: 📅 Thu, Jul 9, 2026 ⏰ 2:30 PM – 4:15 PM KST 📍 HALL A #4411 My collaborators @Sakiazusaaa and @Chufan_Shi will be presenting our work! 🚀
Unified multimodal models can both answer questions and generate images. If we edit one so it says “an apple is blue,” will it also draw a blue apple? Introducing UNIKE — #ICML2026 — the first benchmark for cross-modal knowledge editing in unified multimodal models.
1
4
6
567
Taylor Berg-Kirkpatrick retweeted
Parcae has been accepted to COLM! Feel free to reach out if you want to chat in SF!
We’ve been thinking a lot about scaling laws, wondering if there is a more effective way to scale FLOPs without increasing parameters. Turns out the answer is YES – by looping blocks of layers during training. We find that predictable scaling laws exist for layer looping, allowing us to use looping to achieve the quality of a Transformer twice the size. Our scaling laws suggest that for a fixed parameter budget, data and looping should be increased in tandem! 🧵👇
4
6
33
4,481