VP, Kernels @togethercompute Assistant Professor @ucsd_cse Looking for talented kernel engineers and performance engineers!

Excited to share that I will be joining UCSD CSE as an assistant professor in January 2026! I'll be recruiting PhD students from the 2024 application pool - if you're interested in anything ML Sys/efficiency/etc please reach out & put my name on your application! Until then I'll be finishing up some requirements at Stanford (long story...) and hanging out at @togethercompute. Stay tuned for more!
48
40
579
117,753
Super excited to get early access to Vera Rubin to write some kernels and play with the new tensor cores! Enjoy these astronomer thunderkittens looking into the stars and check out the write up of new features below⚡️🐱💫
thunderkittens is now running on nvidia vera rubin! our kernels team got early access to nvl72 and spent the past few days digging through the new isa and bringing nvfp4 + fp8 gemms to life after reworking the kernels for rubin, we pushed them past 22 and 12 pflops respectively — competitive with cublas + cute dsl
3
2
23
2,340
thunderkittens is now running on nvidia vera rubin! our kernels team got early access to nvl72 and spent the past few days digging through the new isa and bringing nvfp4 + fp8 gemms to life after reworking the kernels for rubin, we pushed them past 22 and 12 pflops respectively — competitive with cublas + cute dsl
13
19
137
24,019
Really cool to see folks train looped models at scale! @hayden_prairie has some great work here on the fundamental architecture choices in looped models and how to make them stable (+ higher quality). More soon on the inference-time memory implications :)
It's exciting to see the rumors that Astra is a looped model. For those interested in the math on how to stably train them, I would check out our work Parcae. We also show that looping follows predictable scaling laws! arxiv.org/pdf/2604.12946
2
1
25
4,359
It's exciting to see the rumors that Astra is a looped model. For those interested in the math on how to stably train them, I would check out our work Parcae. We also show that looping follows predictable scaling laws! arxiv.org/pdf/2604.12946
new: OpenAI & others quietly using loop transformers that don't show their 'thinking' when scaled up a leap forward on performance, but sparking concerns inside & outside OpenAI re: security as this takes off
3
11
53
10,457
This is a really cool demo I've wanted to see for a while - cool to see the models + systems develop to really make a compelling demo ("realtime" video has been possible since 2024 with OpenSora on 8xH100, but quality was not good enough to be compelling) Great work @haoailab!
Wow, the community is moving so fast. @reactorworld just developed an infinite video stream on Twitch using FastH3, and it looks so fun! It is open source!! Thank you @reactorworld for the love with OSS 🥰🥳 code and models here: github.com/reactor-team/infi… haoailab.com/blogs/fasth3-pr…
2
1
18
4,451
♉️𝛂🐂𝛂♉️𝛂🐂𝛂♉️𝛂🐂𝛂
Excited to share that 0x Alpha will be on @togethercompute as early as tomorrow morning! together.ai/models/ox-alpha
1
8
3,559
Excited to lead the way on quality serving for K3!
.@Kimi_Moonshot benchmarked K3 endpoints across major inference providers. Together AI leads or ties for #1 on 3 of 4 benchmarks: OCRBench, MMMU Pro Vision, and DeepSWE. Open models like Kimi K3 have gotten genuinely good. Frontier-class good. Which makes serving them well essential. Proud of our Research team for holding the bar this high, and proud to see it verified by Moonshot. Run Kimi K3 on Together AI: togetherai.link/k3-x
1
1
27
5,806
.@Kimi_Moonshot benchmarked K3 endpoints across major inference providers. Together AI leads or ties for #1 on 3 of 4 benchmarks: OCRBench, MMMU Pro Vision, and DeepSWE. Open models like Kimi K3 have gotten genuinely good. Frontier-class good. Which makes serving them well essential. Proud of our Research team for holding the bar this high, and proud to see it verified by Moonshot. Run Kimi K3 on Together AI: togetherai.link/k3-x
11
10
55
14,759
Great stuff from @stuart_sul! Software is dead, really excited by what the new world brings.
(1/6) If a prompt can generate every layer of the stack, what are programming abstractions worth? GPU programming is at a weird inflection point. For decades we stacked layers of abstraction to ease our cognitive load. This year, we find ourselves constantly deleting them.
1
9
2,411
Dan Fu retweeted
(1/6) If a prompt can generate every layer of the stack, what are programming abstractions worth? GPU programming is at a weird inflection point. For decades we stacked layers of abstraction to ease our cognitive load. This year, we find ourselves constantly deleting them.
Made with AI
9
32
186
18,429
It has been clear to many of us, and now it’s becoming clear more tangibly, that AI models will commoditize to various degrees. This is of course a difficult business reality if your core business depends on exclusivity on intelligence. But commodity markets are not communism. They are the largest markets on earth. Oil, grain, steel, electricity, memory: trillions clear through them every year, priced by competition among thousands of suppliers. In economic terms, communism is one provider and no price. A commodity market is the precise inverse. The world that actually resembles central planning is the one Dean argues for: a set of protected incumbents, access gated by the state, agencies instructed to manufacture FUD until every regulated buyer, and transitively every tool maker upstream, backs away from cheaper competitors. Open weights don't deter capex. They move it. When the model layer commoditizes, spend shifts to inference, data, tooling, and applications, and builds far broader industrial infrastructure rather than concentrating capital in a handful of companies. Most of our digital infrastructure today, hyperscalers included, runs on open source. The businesses built atop it keep excellent margins and compound at extraordinary rates. Open-weights intelligence will likely rank among the most important economic accelerations in history. It won't be kind to every early incumbent, Linux wasn't kind to Sun Microsystems, but it will be very good for almost everyone else. I suspect OpenAI and Anthropic, given their positions, excellent products, resources and talent density, will be just fine. They will simply hold a little less pricing power. The security theater around Mythos continues to do damage. Of course, there is no evidence for the hysterical claims. The evidence is in fact so thin that proponents of AI's existential risks now openly recommend FUD as the strategy. That should be telling.
Some observations on Kimi: 1. It's a very good model! I don't think its performance can be explained away by distillation or anything like that. In agentic coding sessions, it seems pretty much on par with the best public models of Q1 2026. In my fairly limited use, it also seemed very token hungry. It's not obvious to me that this model is actually that cheap to run. 2. I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks. To be clear, I *myself* might be fine with models presenting this level of marginal risk being open weight, but I am surprised that China is fine with it. I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). The other 25% or so is their lack of compute for customer inference (making China's open-weight strategy an unintended byproduct of US export controls) and the normal Chinese strategy of aggressive exports. For the companies, as opposed to the government, the decision to open source is partially ideological and partially because they are behind, and they know that very few people would pay for sub-frontier models from China. 3. Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. I suspect the reason they are is that they know open-weight models are effectively ungovernable, and they simply like the overall cloak of ungovernability open-weight models create over the whole of AI. It's not a bad strategy; it reminds me of James Scott's recounting of the hill people in "the art of not being governed." Still, in the end, open-weight models deter further AI capex. 4. One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end. You'd be surprised how many 'accelerationists' lobbied me, while I was in government, to support an eleven or twelve-figure federally funded data center so that startups could train models at a subsidy and then give them away for free. There was no other way for AI to progress, they said. Perhaps this is the logical end state of things. Nonetheless, I find myself surprised to see supposed accelerationists excited about such an outcome. I think many of them just don't know what they're doing. Many accelerationists do not view the creation and serving of frontier models as a legitimate business. 5. I would guess that the Trump Administration will at some point realize that their best strategy here would be to create large amounts of regulatory risk around the use of open-weight Chinese models. You don't need to "ban open source" (one of the dumber motifs of AI policy discussion). You just need to direct every agency to issue soft law that creates FUD. "A Federal Reserve Advisory Bulletin found that there may be backdoors in Chinese AI models." It needn't be that well justified. You just create enough regulatory risk that every regulated enterprise backs off. You probably don't want to create so much regulatory risk that you scare off the hyperscalers from serving Chinese models; this will just drive startups to sketchier providers. There's a happy middle ground here. I'd assume they will do some version of this. 6. It's probably true that open-weight models of this capability make the world a bit more dangerous, but not so much more that you'll really notice. At some point the models will be capable enough that you will notice. "A nonliving, invisible, dangerous, and infinitely self-replicating agent escaped from a Chinese lab," you say? Color me shocked.
30
72
413
67,601
🎉 Thrilled to share we had 4 papers accepted to COLM 2026! Huge congrats to Hayden Prairie @hayden_prairie, Ivan Lee @ivn1e (lead on two!), Yasaman Jafari @yasjafarii , Zixian Wang, Cheng Yang @ChengYANG_yc , and Zachary Novack @zacknovack — 🧵 below with all the papers: 1/ "Parcae: Scaling Laws For Stable Looped Language Models" Hayden Prairie, w/ Zachary Novack, in collab with Dan Fu @realDanFu Viewing layer looping as a dynamical system yields stable looped LMs, and a new predictable scaling axis at constant memory. 📄 arxiv.org/abs/2604.12946 2/ "Studying the Soupability of Documents in State Space Models" Yasaman Jafari, w/ Zixian Wang & Leon Bergen Encode docs separately with Mamba2, then average the SSM states into one "soup": better multi-doc QA at a fraction of the inference cost, scaling to 256 docs. 📄 arxiv.org/abs/2505.24033 3/ "The Format Tax: Measuring and Mitigating the Cost of Structured Output" Ivan Lee, w/ Loris D'Antoni Forcing JSON/XML output hurts reasoning. The culprit is prompt-side distribution shift, not constrained decoding. Reason freely, then reformat: most of the accuracy comes back. 📄 arxiv.org/abs/2604.03616 4/ "Optical Context Compression Is Just (Bad) Autoencoding" Ivan Lee, w/ Cheng Yang Vision tokens aren't magic: simple baselines like mean pooling match DeepSeek-OCR-style compression on reconstruction and beat it for language modeling. 📄 arxiv.org/abs/2512.03643 #COLM2026
2
11
70
7,326
Dan Fu retweeted
So excited to share that I've recently defended my PhD and joined @ColumbiaLaw as an Associate Professor! I'm absolutely ecstatic to be joining such an incredible community, and looking forward to much future work at the intersection of law and AI. Doing a JD/PhD in computer science at @StanfordLaw and with @HazyResearch was the experience of a lifetime, and I'm so grateful for the opportunity. If you're a student interested in this intersection–please reach out!!
26
23
243
26,075
Our CEO @vipulved on @CNBC with @dee_bosa: your data is your recipe. As models get smarter, sending proprietary workflows, customer context, and business logic into closed systems becomes a strategic decision. Open models help companies build AI while keeping more of their intelligence layer under their own control.
4
9
44
5,325
Dan Fu retweeted
We @togethercompute believe intelligence should be abundant, not expensive. Today we announced our Series C funding of $800m @ $8.3B valuation, to continue to build the world's most efficient platform for generative AI. Thanks @nikogallogly for telling our story in @nytimes! shorturl.at/SooOP
67
87
487
174,921
Multi-GPU kernels are the real test for coding models. Today at @aiDotEngineer, @simran_s_arora shared ParallelKernelBench, an open-source benchmark for evaluating whether LLMs can write fast CUDA kernels for real communication-heavy workloads. Proud to see this work from the Together AI Frontier Performance team.
8
30
246
17,126
Our research team has 9 papers at ICML next week! Spanning the full stack from frontier agents to GPU kernels, we're excited to share what our researchers and collaborators have been working on. If you're in Seoul for ICML come meet our team or catch a talk! 🧵
5
11
75
4,740
Excited to chat with @olive_jy_song live next week on stage at @aiDotEngineer about MiniMax 3! It’ll be a fun one, come check it out :)
next week at @aiDotEngineer, we are joining @togethercompute for a conversation on what goes into running agents at scale. @olive_jy_song, Research Lead, RL at MiniMax, and @realDanFu, VP of Kernels at Together AI, will walk through both sides of M3: the training decisions behind long-context reasoning and tool use, and the infrastructure work needed to serve agentic workloads in production. Wed July 1, 2:50 PM PDT Room 2016, Track 9 Catch you there!
5
4
42
9,374