Founder @RicursiveAI, Asst. Prof. of CS at Stanford. Prev: DeepMind, Anthropic, Brain. Co-Creator of MoEs, AlphaChip, Test Time Scaling.

Stanford, CA
Thrilled to share that @annadgoldie and I are launching @RicursiveAI, a frontier lab enabling recursive self-improvement through AIs that design their own chips. Our vision for transforming chip design began with AlphaChip, an AI for layout optimization used to design four generations of TPUs, data center CPUs, and smartphones. AlphaChip offered a glimpse into a future where AI designs the silicon that fuels it. Ricursive extends this vision to the entire chip stack, building AI that architects, verifies, and implements silicon, enabling models and chips to co-evolve in a tight loop. We sat down with WSJ’s @berber_jin1 to discuss Ricursive: wsj.com/tech/this-ai-startup…
Introducing Ricursive Intelligence, a frontier AI lab enabling a recursive self-improvement loop between AI and the chips that fuel it. Learn more at ricursive.com
127
133
1,536
258,691
CLM-8B vs. Jev
Replying to @jackyk02
🧵(5) Zero-Shot Evaluation Across computer-use, gaming, and tool-calling tasks, CLM-8B performs on par with Jev while running up to 9× faster. The speedups are most pronounced when the number of candidates is large (e.g., WikiRacing) or when actions can be reused frequently across states (e.g., T-Rex Game). CLM-35B, with improved generalization and even greater speedups, will be released early next month.
3
2
63
6,082
Introducing Contrastive Language Models (CLMs), a System 1 model that connects actions and states! Agentic coding: with lightweight finetuning, CLM-8B sets a new SOTA on DeepSWE (81.6%) and Terminal Bench 2.1 (87.6%). CLM-8B vs. Jev: comparable zero-shot performance on computer-use, gaming, and tool-calling tasks, while being up to 9× faster. We are releasing the CLM-8B checkpoint, its data, and infra today!
Introducing Contrastive Language Model (CLM): an ultra-fast System One Model trained with a contrastive learning objective that connects states and actions. CLM-8B is pre-trained on internet-scale data and delivers up to 9× faster inference than Jev ⚡ while achieving comparable performance across computer-use, gaming, and tool-calling tasks. With lightweight fine-tuning, CLM-8B sets a new SOTA on challenging agentic coding benchmarks, such as DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%). In contrast, Jev fails to serve as an effective verifier for these long-horizon tasks. We also build an efficient training and serving infra for CLMs by disaggregating states and actions, allowing their embeddings to be cached and reused independently. This substantially reduces inference latency in settings where the state evolves continuously while the action set remains fixed. Finally, we establish scaling laws for CLMs and show that the test contrastive loss decreases predictably as a power law in training compute, model size, and dataset size. 📄 Blog: contrastive-lm.notion.site 💻 Code: github.com/Contrastive-LM/CL… 🗣️ Discord: discord.gg/5dAQEDJBs 🤗 Data & Models: huggingface.co/Contrastive-L… More details on CLM’s architecture, data recipe, and scaling laws in the thread below 🧵
46
100
1,299
104,300
Azalia Mirhoseini retweeted
1/ some personal news! I’m back in chips and joined Ricursive Intelligence as Head of Partnerships and Business Development to work with @AnnadGoldie & @Azaliamirh and make deals with chipmakers, fabs, and AI companies to accelerate chip development, provide fundamental infrastructure for model-hardware-software codesign, and create a specialized recursive research loop for hardware.
53
5
227
34,484
Welcome to the team, Michelle! Super excited to work together!
1/ some personal news! I’m back in chips and joined Ricursive Intelligence as Head of Partnerships and Business Development to work with @AnnadGoldie & @Azaliamirh and make deals with chipmakers, fabs, and AI companies to accelerate chip development, provide fundamental infrastructure for model-hardware-software codesign, and create a specialized recursive research loop for hardware.
4
1
57
8,337
Check out Turbo-dLLM, our new open-source library for training diffusion LLMs at scale! For example, on 8x H100 GPUs, Turbo-dLLM accelerates DFlash2 speculative-decoder training by 2.48x at 512K and 7.59x at 1M context length. Turbo-dLLM introduces Context-Sharded Block Parallelism, a new paralleism strategy that unlocks significant efficiency gains for training, especially as the context length increases. Great work led by @TarunSures41845, @PranshuChatur11, @hangoo_kang and an amazing team! Github: github.com/ScalingIntelligen… Website: scalingintelligence.stanford…
Diffusion LLMs and speculative decoding promise much faster agents. Yet agents need long contexts, and training on them is painfully slow. Introducing Context-Sharded Block Parallelism (CSBP), a new distributed parallelism strategy unlocking significant training efficiency for diffusion LLMs, with speedup gains growing with context length 🚀 ⚡ 7.59× faster DFlash2 speculative decoding drafter training ⚡ 1.61× faster block diffusion fine-tuning ⚡ 1.33× faster autoregressive → block diffusion adaptation With the same GPU hours, models trained with CSBP score higher on SWE-bench Verified and Terminal-Bench Lite 👑 Open-sourced in Turbo-dLLM, our new optimized distributed training library. Advised by @Azaliamirh and with an amazing team: @PranshuChatur11 @hangoo_kang @pshroff_ @ishanskhare @KumbongHermann
6
22
150
12,711
The inference landscape is going to get a lot more hybrid in the near future. We found that accuracy per joule of local models has improved 18x in just 16 months: 5.9x from hardware, 3.0x from model gains. Great in-depth cover by @FT: ft.trib.al/eVI82ZQ @Avanika15 @JonSaadFalcon John Hennessy @HazyResearch
dreams do come true 🥹. excited to see our work (w/@JonSaadFalcon, @HazyResearch, john hennessy and @Azaliamirh) feat. in a major way in @FT. the world is becoming increasingly less dependent on centralized cloud ai. we are just getting started 🚀🌖
12
36
253
30,011
Azalia Mirhoseini retweeted
From robot learning to chip design to AI security, Stanford faculty @sanmikoyejo, @chelseabfinn, and @Azaliamirh are shaping where AI goes next. Congratulations to these three on being named to the TIME AI 100! ti.me/100ai
3
8
56
4,443
Azalia Mirhoseini retweeted
We at @Felicis are true believers in @RicursiveAI. Congrats @Azaliamirh and @annadgoldie on this impressive honor.
Honored to be named to @TIME's TIME100 List of the World's Most Influential People in AI, even more so to share it with my co-founder @annadgoldie! Anna and I started working on AI for Chip Design almost a decade ago. Last year, we started @RicursiveAI to transform end-to-end chip design from years to days! Watching that vision become reality piece by piece has been the most thrilling / fulfilling experience ever! time.com/collection/time100-…
2
1
5
1,704
Thanks for being an amazing partner along the way, Stephanie!
So well deserved! The AI era is compute-constrained, and chip design is one of the deepest bottlenecks. @annadgoldie & @Azaliamirh are true forces of nature: they saw this early (before this has now become obvious!) and built @RicursiveAI to solve for it Now they’re proving AI for chip design in production: real industrial chip designs, against commercial tools, with step-function results: - Dramatically faster runs - Cleaner layouts - The ability to take on problems existing workflows struggle to handle Their early results are groundbreaking, and everyone from chip incumbents, frontier labs, to hedge funds are now asking for it
2
19
5,295
Azalia Mirhoseini retweeted
These women are absolute rockstars! If you are interested in joining a fast growing AI Chip Design company or are passionate about semis check out @RicursiveAI!
Honored to be named to @TIME's TIME100 List of the World's Most Influential People in AI, even more so to share it with my co-founder @annadgoldie! Anna and I started working on AI for Chip Design almost a decade ago. Last year, we started @RicursiveAI to transform end-to-end chip design from years to days! Watching that vision become reality piece by piece has been the most thrilling / fulfilling experience ever! time.com/collection/time100-…
2
1
13
1,864
Honored to be named to @TIME's TIME100 List of the World's Most Influential People in AI, even more so to share it with my co-founder @annadgoldie! Anna and I started working on AI for Chip Design almost a decade ago. Last year, we started @RicursiveAI to transform end-to-end chip design from years to days! Watching that vision become reality piece by piece has been the most thrilling / fulfilling experience ever! time.com/collection/time100-…
33
23
219
15,914
Check out Hawkeye, which writes high-performance kernels utilizing advanced architectural features (e.g., TMA for async data transfer on Blackwell, L2 locality on MI350) from only one handwritten example and ~10 unit tests. Hawkeye can port kernels across architectures (Ampere, Hopper, Blackwell), chips (NVIDIA, AMD), and precisions (FP8, NVFP4, MXFP4). Great work, co-led by @AryaTschand and @keramakr!
We’ve seen an explosion of new ML chips with unique architectural features, but software support remains the critical bottleneck Achieving peak performance increasingly relies on hardware-specific optimizations in the kernels, but we observe that coding agents are particularly weak at this Introducing Hawkeye, a framework that brings hardware-awareness to coding agents by grounding them in a minimal and comprehensive taxonomy of optimization strategies For new GPU or ML accelerator architectures, you only need to write 10 unit tests and solution kernels (one per optimization strategy), and we show that coding agents can effectively scale test-time compute with this minimal supervision to write hardware-aware kernels Hawkeye can port kernels across architectures (Ampere, Hopper, Blackwell), vendors (NVIDIA, AMD), and precisions (FP8, NVFP4, MXFP4) while consistently leveraging hardware features and approaching expert kernel performance Work co-led with @keramakr and done in collaboration with Alexander Ingare @simonguozirui @18jeffreyma @ZishenW @simran_s_arora @Azaliamirh @profvjreddi
2
17
208
17,989
📈
Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper 💰 As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost. For example, we find that sampling just 5 solutions with DeepSeek V4 Flash and ranking them using the same model with LLM-as-a-Verifier can lead to a significant boost in accuracy (79% → 88%), outperforming closed frontier models on Terminal-Bench. Try it out today: github.com/llm-as-a-verifier… More on verification scaling in my previous post.
1
39
5,918
LLM-as-a-Verifier keeps pushing the frontier of cost vs. capability! On Terminal-Bench 2.1, it made DeepSeek V4 Flash accuracy go from 79% → 88%, while being 4-11x cheaper than competitors! Try it here: github.com/llm-as-a-verifier… @jackyk02
Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper 💰 As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost. For example, we find that sampling just 5 solutions with DeepSeek V4 Flash and ranking them using the same model with LLM-as-a-Verifier can lead to a significant boost in accuracy (79% → 88%), outperforming closed frontier models on Terminal-Bench. Try it out today: github.com/llm-as-a-verifier… More on verification scaling in my previous post.
20
50
595
57,686
Thanks, Jeff! So excited for you all and rooting for you! ❤️
Thanks, @Azaliamirh, and thanks to both you and @annadgoldie for being so kind in offering advice to us as we start this journey and having us over for Ricursive's happy hour a little while ago! We're inspired by what you all are building there!
6
1
120
27,868
Congratulations to @JeffDean, @Sanjay_Ghemawat, @quocleix, and @OriolVinyalsML on Discovery Loop!!
Announcing Discovery Loop! I am very excited to announce that, along with my longtime friends and collaborators @Sanjay_Ghemawat, @OriolVinyalsML and @quocleix, we are founding Discovery Loop (@DiscoLoopAI), a Public Benefit Corporation whose mission is to automate machine learning, science, and engineering to accelerate discoveries and progress. The four of us have worked together for 14 to 30 years, and have helped build some of the world’s most used products, infrastructure and AI models, and we’re excited to turn our attention to this ambitious endeavor. ♾ Learn more at: discoveryloop.com
5
4
144
16,717
I had so much fun creating the Self-Improving AI Agents course with @achowdhery, and teaching it twice in one year at Stanford! We also collaborated with Stanford Online to make the course available online: YouTube: piped.video/6YnLB0XbTnI?si=MVwR… The field is moving incredibly fast, but we tried to focus on the core concepts that help us build better AI systems. I hope you enjoy it!
41
119
1,144
106,884
Azalia Mirhoseini retweeted
"AI inference demand is expected to grow 10,000x over the next 5 years." Yesterday at @DACconference, @annadgoldie took the stage to share how we break today's deadlock between model and hardware. That's what we're doing at Ricursive, using our very own model. 📍 Stop by Booth #1049 to say hi.
9
21
4,612