assistant prof @mldcmu. chief scientist @cartesia. leading the ssm revolution.

🤔
Albert Gu aka the Mamba guy has the top ranked voice model now. How much of that is the architecture and how much is they just trained it better?
1
52
10,160
chat with your model with a single identity across languages and accents!
Introducing Multilingual Voices We partnered with Baklavastory, our neighbor in the Mission, to show what it sounds like when a business's character and warmth stay consistent, no matter who calls or what language they speak. It's easy to translate speech, but much harder to preserve a single identity across languages. Every language has different rhythms, tones, and emphasis. When you change the language, identity tends to get lost in that shift, and your voice ends up sounding like someone else. With Multilingual Voices, pick a voice and keep one brand identity in every market you serve: cartesia.ai/voices
33
5,095
poor ronald 😭
We regret to inform our AI researcher Ronald that AI has replaced him. Sorry Ronald. Here's our team trying to tell his real voice from his cloned one. They were not good at it.
42
10,334
misleading paper title, and even the phrasing in the tweet is still a misnomer. this is not about "post-training" but "retrofitting" (cross-architecture distillation) fwiw i've been bearish on distilling Transformers to recurrent models for a while, they're too different; architectural innovations just need to be trained from scratch
Simple beats complicated: We show that switching to a sliding-window attention mask with attention sinks (at no cost) beats linear attention post-training. Huge thanks to my collaborators @RheaSukthanker, @CameronPashmina, and @Emy_Aze. Paper: arxiv.org/abs/2608.28444
16
24
656
80,556
Albert Gu retweeted
Cartesia Sonic-3.6 is the #1 TTS model on the Voice Arena US English 🇺🇸 Leaderboard. It debuts at 1085 Elo, in a statistical tie with Google DeepMind's Gemini 3.1 Flash TTS, and ahead of its own Sonic-3.5, Simba 3.2 and Grok TTS. It holds the top spot as a real-time model. Sonic-3.6 is the latest TTS model from @cartesia. The model has been highly preferred among raters on @voicearena_ai. Results backed by 1,400+ blind, head-to-head listener votes across two languages 🧵
7
101
145
664,217
sonic-3.6 is out, with a large improvement over (the already #1) sonic-3.5 in just a few months! the research team's focus on fundamentals is accelerating progress at the frontier of architectures and audio
Sonic-3.6 is now generally available. In January we made a bet: stop tuning the existing paradigm, rebuild from the architecture up. How we topped our own best model in two months → cartesia.ai/blog/sonic-3.6/?…
3
2
85
13,997
Albert Gu retweeted
Introducing Sonic-3.6: our most lifelike TTS yet, and another step change in naturalness across 44 languages. Just three months after our Sonic-3.5 launch, we've made fundamental model improvements based on feedback from the teams building on Sonic. We're #1 on @ArtificialAnlys across both the provider and controlled voice streaming leaderboards — best-in-class voices + the best underlying model. Available in beta today. Hear it for yourself 👇
39
73
664
90,300
the team continues cooking 👩‍🍳 this is now an unprecedented gap on both the super competitive Provider Voices leaderboard as well as the newer Controlled Voices leaderboard (a stronger benchmark that can't be benchmaxxed, requiring truly better algorithms) Cartesia's TTS model is not only the best, it continues improving at a faster rate than the competition - Sonic 3.5 was released only 3 months ago. ig Sonic now means both the fastest model and the fastest team 💀
Cartesia's Sonic 3.6 takes the #1 spot on both the Provider Voice and Controlled Voice Artificial Analysis Speech Arena leaderboards, surpassing Speechify AI's Simba 3.2 and Alibaba's Qwen-Audio-3.0-TTS-Plus, with Sonic 3.5 holding #2 on Controlled Voice Sonic 3.6 is the latest TTS model from @cartesia, supporting 40+ languages including English, Hindi, Spanish, French, German and Japanese. In the TTS Arena, the model has been preferred by voters and has demonstrated natural speech and adaptive pacing across conversations. Key takeaways: ➤ Controlled Voice: Sonic 3.6 takes #1 on the Controlled Voice Arena with an Elo of 1,144 (+18/-18) across 1,339 appearances, ahead of Cartesia's own Sonic 3.5 at 1,099 and ElevenLabs' Eleven v3 at 1,060 ➤ Provider Voice: Sonic 3.6 also takes #1 on the Provider Voice Arena with an Elo score of 1,286 (+19/-19) based on 1,300 arena appearances, placing it ahead of Simba 3.2 at 1,238 and Qwen-Audio-3.0-TTS-Plus at 1,237 ➤ Pricing: At $49 per 1M characters via the Cartesia platform, Sonic 3.6 is the most expensive of the top three models on quality, roughly 5x Speechify AI's Simba 3.2 at $10 and nearly double Alibaba's Qwen-Audio-3.0-TTS-Plus at $27.59, though still half the price of ElevenLabs' Eleven v3 at $100 ➤ Speed: Sonic 3.6 processes 136.1 characters per second of generation time, compared to 46.7 for ElevenLabs' Eleven v3 and 24.2 for Google's Gemini 3.1 Flash TTS
5
10
111
12,376
Albert Gu retweeted
A very happy 80th Independence Day to India, from team Cartesia! 🇮🇳 Coming closer to home with Sonic 3.6, we're making every one of India's languages more natural, expressive, and accurate: - Hindi that flows the way you speak it - Hinglish that switches without missing a beat - Naturalness gains across all 9 existing Indic languages - Plus two new languages: Urdu and Odia Every voice in this video generated by Sonic.
90
205
508
111,630
Albert Gu retweeted
Excited that @cartesia is in @cursor_ai's India campaign with some of the most ambitious founders of Indian origin (including in Delhi where I grew up) featuring amazing folks @amanrsanger, @vipulved, @tankots, @mukundjha, @ManishaRaisingh, @regards_rishi, @_sankyy
4
5
109
7,272
Albert Gu retweeted
Our London office is growing quickly, reach out to @_albertgu or me if you’d like to work on a very different, new and exciting research agenda to advance multimodal models
Good night DeepMind. Wow. FT: "Google is shifting control of its AI effort from London back to Silicon Valley" "executives and board members remained concerned by its weaker position in coding models and enterprise AI, where Anthropic and OpenAI have established an early lead." "Several current and former DeepMind employees said the reorganisation had sent shockwaves through the London-based lab, with some fearing it marked the end of the research culture that Hassabis has long protected. One former DeepMind executive at a rival company said they had already received calls from multiple staff who said they were ready to jump ship." "Over the past year, Hassabis has devoted more time to Isomorphic Labs, the AI drug discovery company he founded"
1
4
88
17,531
Evaluations are difficult and vague for all generative models, and benchmarks only capture a small slice. Our blog post dives into the nuances for TTS
"Is this TTS model good?" gets harder to answer as models improve. "Good" is at least five axes: correctness, naturalness, contextual correctness, robustness, and most evals only capture the first. We wrote up the failure modes that make TTS eval hard: cartesia.ai/blog/is-this-tts…
3
4
54
12,296
Albert Gu retweeted
We will be at ICML in Seoul to present dnaHNet in Oral + Poster sessions! Swing by and say hi if you wanna learn about the current SOTA in DNA foundation models :) Oral session: Hall D2 at 10 AM, Jul 7 Poster session: Hall A #900 at 2 PM, Jul 7
Most genomic AI models use fixed rules to process DNA into chunks, imposing arbitrary boundaries on a sequence with its own biological structure. @arnavshah0, @victor_ljz, and team developed dnaHNet, a tokenizer-free foundation model that learns its own segmentation from scratch, supervised by @_albertgu, @genophoria, and @BoWang87.
7
17
60
21,950
Transformers are better at copying, while RNNs are better at modeling "meaning-bearing words—the nouns, verbs, & adjectives that say what a sentence is about"
Hybrid (transformer–RNN) models are fast becoming a serious alternative to the transformer, but a big question remains: how do they process tokens differently & how does this impact performance? We compared our transformer (Olmo 3) & hybrid (Olmo Hybrid) models to find out. 🧵
6
29
428
60,437
Rather than interleaving layers naively, a more fine-grained approach to hybrid models is to allow hybridization across the sequence models within a single layer. The fact that softmax attention and linear attention use similar underlying projection parameters allows switching between different mixers in a single generation, for the best of both worlds.
Excited to share last summer's work at Google Research! Most hybrid models today are static: each token sees the same interleaved pattern of your favorite linear model and attention. Oryx instead varies the model used across the sequence through shared representations. 1/
4
17
229
33,119
Congrats to Henry and Naomi - they’ve been so on top of the space and super helpful as collaborators too!
Most AI investing happens downstream of the frontier: a capability emerges, a category gets named, and capital rushes in. But by the time a category earns a clean box on a market map, the best builders have usually been living in the messy version for months. Agents. Reasoning. RL environments. World models. AI for Science. Recursive self-improvement. I call this frontier proximity: the ability to see what is becoming possible before it becomes consensus. My frontier proximity ladder: L0 Wrapper: uses today’s models. L1 Reactor: reacts fast to releases, but roadmap is downstream. L2 Anticipator: builds for where capabilities are going. L3 Native: depends on a non-obvious frontier bet. L4 Shaper: helps move the frontier itself. The point is not that every company needs to train models. Apps can have high frontier proximity if they understand what models will make possible next. Infra can have high frontier proximity if it knows what future agents, multimodal systems, robotics stacks, or scientific workflows will need. That is why we’re launching MoE Capital. MoE stands for Mixture of Experts. The idea is simple: build an AI fund around people closest to the frontier: frontier researchers, technical founders, AI-native builders, and seasoned operators. We don’t want to be another AI fund with a newsletter-level understanding of the frontier. We want to build the AI fund closest to the frontier. More in The Information: theinformation.com/newslette…
4
1
29
9,857
Albert Gu retweeted
Two new models just dropped 👀 Sonic-3.5 and Ink-2 are the #1 streaming models for text to speech and speech to text
We released Sonic-3.5 and Ink-2, the #1 streaming models for text to speech and speech to text you can use in your voice agents today. New architectures enable new frontiers for speed and quality. We're now the only provider to have #1 models for both speaking and listening.
14
37
147
30,196