🧑‍💻 Voice AI @trydaily // ᓚᘏᗢ @pipecat_ai 🎹 I have too many synths

London, UK
Trying out Opus 5.5, asked it to one-shot an infinite stickman duel (I miss Flash 😞) Watch them duke it out here, forever! Or until I run out of Digital Ocean credits: twitch.tv/makeitmoreai Incredibly, it even added random hazards, weapon drops, pretty decent fighter AI, physics, different arenas, super saiyan mode, berserker mode, bullet-time KO cam, and probably even more things I haven't seen yet. It even thought to generate sound effects 🤯 Go blue fighter!
6
447
(warning: tweet not Jev related 😅) Some updates to Pipecat Visemes, a drop-in DSP lipsync processor for @pipecat_ai bots: playout-timed mouth keyframes from any TTS, for client-rendered avatars. ➡️ Server-side lipsync for any Pipecat TTS. ➡️ numpy-only, usage only ~1% of a core per bot. ➡️ Delivery timing fixed: the mouth now anchors within ~3 ms of the audio at utterance start (was +300 ms), with keyframes sent 200 ms ahead of playout. ➡️ 87.5 / 90.7 on our Praat-scored benchmark. Code: github.com/jptaylor/pipecat-…
1
11
952
Wow, @typesafeai Jev is a game changer for voice-driven UI. Responsiveness really matters for an experience to 'feel' right, and this certainly does the job! Latency here is not optimal given I'm in London and calling out to us-west. But still, super snappy. In this video, Jamcat is using Pipecat to orchestrate a real-time handover between a command agent (Jev) that drives the UI, and voice agent (PhoneLLM) for handling generations and (not in the video) discussing ideas and session setup with the user.
18
37
479
47,252
Audition mode is coming along in Jamcat. The generated midi guide is used in audio-to-audio to generate stems that match your voice prompt. Each instrument type needs tuning and training, but it's not far off!
1
1
6
335
Jamcat's song starter mode is addictive! Check out this rock jam, made entirely by voice prompting (I didn't do any vocal auditioning in this video, but it's possible.) Offering a simpler UI with no clickable elements is a fun way to get a groove going before switching into DAW mode, or exporting to Ableton. Voice-driven UI is hard. You need a good frontend for your STT, and a clear design hierarchy for UX and prompting. Music is a great way to explore these concepts, because you can rock out whilst learning what works and what doesn't.
2
2
5
527
And here we are 🤯 chucking the Ableton Live export over to Astral. Full play-through at the end. It structured the track, added new elements (in harmony!), transition FX, gain staged etc. Incredible to watch it work for 20 minutes.
1
3
174
Over a year ago, I posted a video where Gemini and I collaborated on an Ableton Live Session. Music AI is fun! Creating voice-native music software is something I'm really passionate about, so, here is Jamcat. Jamcat's concept is a voice controlled, generative music performance tool, designed to inspire live, in the moment jam sessions. It uses LLMs, with both text-to-audio and audio-to-audio models to create music in realtime, entirely hands-free. ➡️ Pipecat PhoneLLM Alpha 1 for session control, running on @modal. ➡️ "StemGen" routing and workload subagent. ➡️ Foundation 1, for melodic stems. ➡️ Various audio models on @fal (like the excellent @ElevenLabs Music) for vocals, drums, textures and one-shots. ➡️ ... and of course @pipecat_ai! Voice commands can be toggled via push-to-talk (or a foot pedal, if you're wielding an axe), or it can listen the entire time (diarization.) It makes some mistakes, or sometimes misses a beat, but most of the mistakes in this video were me not paying attention to the bar counts, like a true fake musician.
4
4
38
7,316
Here is an example Pipecat project for trying PhoneLLM Alpha 1. Deploy the model to a Modal endpoint with one click. Client from the video in repo too (with all that sweet sweet terminal-ish aura.) github.com/pipecat-ai/pipeca… Next up, a TTS that can pronounce @bmervetan's name correctly? 😅
6
22
221
21,109
Jon Taylor retweeted
Introducing PhoneLLM, an open model for voice agents. GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost. For voice agents, we need models that are both very low latency and very good at tool calling and instruction following. There's a trade-off here, and we often have to compromise on either latency or capability when building voice agents. With PhoneLLM (and the training and data stack that made this model possible) we're fixing this problem. For the last couple of years, most of the effort in frontier model development has gone towards leveraging test-time compute. Which is awesome! Models of all shapes and sizes are available that perform really, really well ... if you have "thinking" turned on for your model. But if you need your agent to respond at voice conversation speed, you can't use thinking models. PhoneLLM is a full-weights fine-tune of NVIDIA Nemotron Nano 30B. We trained on a wide range of real-world telephone and customer support use cases. The training focused on taking the excellent Nano 30B base capabilities and teaching the model to do typical voice agent tasks with thinking disabled. The results are really good: accurate tool calling and concise, on-topic responses in long conversations. And fast: TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200. :-) But seriously, when we characterize model latency, we do it with full, end-to-end, batched request simulations using real Pipecat voice agent pipelines. You can serve more than 80 concurrent agents on a single B200 with P95 end-to-end TTFAT <600ms. Including network overhead. That's an LLM cost-per-minute around $0.0025. (1/4 of a cent.) At a latency lower than any third-party API offers today. More details about this model, including weights on @huggingface, how to spin it up with one click on Modal, and a starter project repo you can clone, are in the thread ...
119
251
2,536
331,713
Oh, are we animating our 'little guys'? The @pipecat_ai mascot was made for this moment...
1
16
6,637
Jon Taylor retweeted
Congrats to @DeepgramAI on the GA launch of Flux TTS, now live in Pipecat 🎉 @JonPTaylor looks at how Flux TTS delivers a more consistent voice experience. Flux reads the whole conversation, not just the next line — adaptive tone, consistent pronunciation, clean interruption handling natively (no SSML markup, no style tags!) all at sub-200ms latency. ➡️ docs.pipecat.ai/api-referenc…
11
24
163
22,639
Ever measured voice agent performance by gap measuring in Audacity? I know I have. Here is a system-level app that measures turn-by-turn metrics (including DAC/ADC and Bluetooth latency estimates.) There is also a handy WS bridge to overlay Pipecat RTVI events 😽 Code here: github.com/jptaylor/pipecat-…
1
2
6
347
Voice agent UI feels stuck in the wiggly-waveform / glowing-orb phase. High-end avatars are exciting, but simple client-rendered characters need richer data than just amplitude to convey speech. An RMS-driven mouth flap gets you pretty far, but believable mouths need phonetic visemes — which usually means provider-specific viseme/timestamp APIs that only some TTS services offer. pipecat-visemes is a drop-in for Pipecat bots that works with any TTS provider: it runs lightweight formant analysis (LPC, numpy only — no ML) on the TTS audio stream itself and ships mouth keyframes to the client as standard RTVI messages. Keyframes ride the output transport's clock, so they land ~200ms ahead of the matching audio, stay in sync, and vanish on interruption. The browser just renders. Benchmark harness (in repo) scores the analyzer against Praat reference formant tracks and a phonetically-designed corpus ("Father was calm as he walked past the palm trees…"), checking raw formant error, mouth-trajectory shape, and articulation events. Tuning runs re-score offline against a committed baseline, so every DSP tweak shows up as a number. The formant analyzer is Tier 0 of a planned ladder: provider viseme events, timestamp+G2P, and an ONNX phoneme model can all slot in behind the same wire format — clients never change, accuracy just goes up.
3
3
33
7,456
AIEWF next week, joining my fellow Pipecat teammates to talk all things voice agents. We'll be at booth U-G8 with our friends at Gradium. And you better believe there will be snazzy new Pipecat swag locked and loaded 👕
2
2
7
389
Arrived in Madrid. Spicy hot, but pumped for the Pipecat + Deepgram + AWS voice AI networking event this evening. Come say hi! (I’ll be the guy wearing a cat t-shirt, of course) luma.com/voiceai-madrid?tk=e…
1
353
Jon Taylor retweeted
I'm obsessed right now with figuring out the right patterns for LLM subagents. I think that all software we write is going to have a lot of inference loops running all the time! A big driver for this is that human-in-the-loop processes must be fast and non-blocking. (Think about voice interfaces, for example.) These human-in-the-loop LLM ... um, loops ... also need to start, stop, steer, and share context with longer running tasks that require external resources, need to use bigger/slower models, etc. We've been hacking on a bunch of subagent abstraction experiments in Pipecat. One thing we need is good ways to run subagents both locally and remotely. Enter @vercel sandboxes! Gradient Bang is our big, open source, multi-player LLM game. It's a canvas for experimenting with subagents, very long contexts over many sessions, dynamic user interfaces, and voice. @jonptaylor just added "bring your own subagents" support to Gradient Bang, built on Vercel sandboxes. Here's his video walk-through.
6
5
27
1,767
Landing today in Gradient Bang: further combat updates for observed / indirect encounters. Useful when you prefer having corp ships as muscle (keeping your personal ship safe and sound in fed space!) Updating the map and ship status on the UI was a much needed feature.
2
1,982
Claude has figured out how to name plans in such a way that gets my immediate approval without even reading them 🐬✨
1
121
Jon Taylor retweeted
✨ Voice AI, open models, and next-generation evals hackathon at @ycombinator in SF on May 30th. ✨ We're co-hosting with @cekuraAi , and we've pulled in our friends at @NVIDIAAIDev, @AWS, and @twilio for expertise and mentoring. We'll help you build state of the art voice agents using: - NVIDIA Nemotron models - AWS SageMaker and Bedrock inference - Twilio telephony - Cekura evaluation tooling - Pipecat orchestration and Pipecat Cloud agent hosting Up for grabs: - A guaranteed YC interview - Special judges' prizes from NVIDIA, AWS, and Twilio for the most impactful and technically impressive projects Join us to learn from engineers who built all the tools you're using, compare notes with other voice AI developers, and show off your ideas! Space is limited. Apply below.
21
34
249
53,206