AGI will not be typed, it will be spoken: that's the closing line of "Speech-to-Speech Model Research at Google DeepMind," the talk
@valeriawu_ and Tom Ouyang, both of Google DeepMind, gave, and
@aiDotEngineer has posted it on YouTube. It traces how Gemini's natively multimodal pretraining replaced the old pipeline of separate acoustic, pronunciation, and language models, and shows what one speech-to-speech model can do once translation, tool use, and conversation all run through it.
- Live translation. Gemini's live model does real-time speech translation across 70+ languages, preserving each speaker's voice, with streaming quality that holds up against offline systems that get to hear the whole utterance first.
- Three things pulling against each other. The talk frames the model's goals as conversational quality, intelligence, and multimodality: push "thinking" higher and evals show more intelligence, but latency and naturalness take the hit.
- Knowing when to stay quiet. A feature called proactive audio lets the model decide not to respond when it hears background noise or someone else talking, since most real conversations aren't happening in a quiet room.
- Non-English first. Most Gemini users aren't English speakers, so translation and localization work covers all their languages, not just English.
- Two very different demos, one model. A Search Live clip identifies a boucle sofa in Spanish, localized to Spain's Spanish, correctly leaving "mid-century" in English; a Live API demo has the same model handling a roadside assistance call, reading back a car's registration plate and postcode.
- Faces, not just voices. A pilot with Citi shown at Cloud Next adds real-time, multilingual, lip-synced avatars on top of the same model.
I'm working through the published talks from AI Engineer World's Fair sharing summaries and takeaways. Follow for more!