A month ago, AGI House builders got early‑access hands‑on with Inworld’s latest TTS model. It’s now officially GA! Hear all the details in our new podcast episode ↓
@KylanGibbs is the CEO and co-founder of
@inworld , and previously worked as a product manager on early LLM and text-to-speech efforts at Google and DeepMind. After watching that technology land almost exclusively in enterprise use cases, he set out to build the infrastructure that would let consumer-facing AI actually reach everyone.
In this conversation with
@jelares (CTO of AGI House), Kylan breaks down why voice is quickly becoming the default interface for AI, what it actually takes to serve that experience to millions of users, and why Inworld chose to build an end-to-end stack rather than stitching together providers.
Timestamps
00:00 Intro
00:34 Kylan's Background & Founding Inworld
02:10 Inworld's Focus
03:41 Why Voice Is the Most Natural Interface
04:53 Overview of Inworld's New TTS Model
06:15 Natural Language Steering for Voice
08:04 Voice Cloning & Voice Design
08:38 Scaling Voice to Millions of Users
09:35 Personalizing Voice by Use Case
10:53 Cost & Latency at Scale
12:32 Why Latency Feels Like Meaning to Users
12:47 End-to-End vs. Piecing Together Providers
15:52 Performance Benefits of a Unified Stack
17:35 Inworld's Research & Inference Teams
20:20 Build vs. Buy: The Case for Using Inworld
23:33 Why Voice Adoption Is Reaching an Inflection Point
30:11 What Consumer AI Apps Look Like in 2029
36:04 Closing Thoughts