Google just released Gemini 3.8 Flash TTS. It can stream two-speaker audio and lets you control how each line is delivered.
I used it with Agora RTC to build a live AI podcast. Everyone in the room hears the same show, and listeners can jump in and talk to the hosts.
I interrupt them twice in this video. The hosts acknowledge the interruption immediately, Gemini 3.8 Flash TTS streams a two-host answer, and then the podcast picks up where it left off.
It feels less like listening to an episode and more like stepping into the conversation.
Built with Gemini 3.8 Flash TTS, Gemini 3.5 Transcribe Live, and
@AgoraIO RTC + RTM. I'll open-source it later.