Proud of the team for delivering the most natural and expressive speech model I’ve ever used (by far).
One simple insight was that text-to-speech isn’t a “small model” problem anymore - the more frontier the model, the better it understands what it’s saying, the more it will sound like it understands what it’s saying.
This isn’t just text-to-speech, but frontier audio generation, including an endpoint to design and edit voices using text prompts. Try generating voices with any nuanced dialect, accent, or generational speaking style - AI should communicate in the exact style that specific users are most comfortable with. Or fun things like “Vampire meditation guide with Transylvanian accent,” “Swedish ASMR furniture review,” “goblin auctioneer selling rare mythical objects,” “tinny monotone robot selling neural transplant upgrades,” etc. AI doesn’t need to sound human, it should just be maximally understandable (and fun to use!).
Audio is the primordial way humans communicate. As we approach human-level speech capabilities - which has admittedly taken a bit longer than text - I expect it to become the primary way we interact with AI, online in real-time applications and offline in education, entertainment, etc. And everyone should have their own trusted AI they recognize by voice.
We’re launching Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS ⚡️
Our most expressive audio models yet let you create custom voices across 100+ languages or pick from 2,000+ ready-to-use ones. You can direct back-and-forth conversations, guide the delivery line-by-line, and add natural cues like <laughs> or an active listening interjection like |mhm| all while generating hours of consistent, glitch-free audio.
Sounds pretty cool, right? So… how should you use them?
— Gemini 3.8 Flash TTS: Need to design bespoke vocal personas from scratch and with line-by-line level control? This is the model! Built for high-fidelity creative production like gaming, immersive audiobooks, and podcasts.
— Gemini 3.8 Flash-Lite TTS: Want the AI to automatically adjust its tone and pacing on the fly for near real-time voice agents? This is your engine! Built for cost-efficient scale, high-volume dubbing, and bulk audio creation.