Infrastructure and developer tools for real-time voice, video, and AI. @trydaily // ᓚᘏᗢ // @pipecat_ai

San Francisco, CA
Pinned Tweet
Introducing PhoneLLM, an open model for voice agents. GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost. For voice agents, we need models that are both very low latency and very good at tool calling and instruction following. There's a trade-off here, and we often have to compromise on either latency or capability when building voice agents. With PhoneLLM (and the training and data stack that made this model possible) we're fixing this problem. For the last couple of years, most of the effort in frontier model development has gone towards leveraging test-time compute. Which is awesome! Models of all shapes and sizes are available that perform really, really well ... if you have "thinking" turned on for your model. But if you need your agent to respond at voice conversation speed, you can't use thinking models. PhoneLLM is a full-weights fine-tune of NVIDIA Nemotron Nano 30B. We trained on a wide range of real-world telephone and customer support use cases. The training focused on taking the excellent Nano 30B base capabilities and teaching the model to do typical voice agent tasks with thinking disabled. The results are really good: accurate tool calling and concise, on-topic responses in long conversations. And fast: TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200. :-) But seriously, when we characterize model latency, we do it with full, end-to-end, batched request simulations using real Pipecat voice agent pipelines. You can serve more than 80 concurrent agents on a single B200 with P95 end-to-end TTFAT <600ms. Including network overhead. That's an LLM cost-per-minute around $0.0025. (1/4 of a cent.) At a latency lower than any third-party API offers today. More details about this model, including weights on @huggingface, how to spin it up with one click on Modal, and a starter project repo you can clone, are in the thread ...
119
251
2,537
331,894
Lots of stuff happened this week, and I was at Meta Connect on Wednesday and Thursday so I'm behind on everything non-Meta. We listened to @thursdai_pod on the school drop-off run to catch up. :-) I love what Alex is doing with ThursdAI. He has an AI producer that helps manage the show live. And he's hacking on crazy custom post-production and video editing tools. The best way to understand what every model/agent can and can't do is to build things that you'll actually use every day.
Crap 12 minutes late, but this week DESERVES a late @thursdai_pod arrival! Opus 5.5 is the GOAT model - the best AI model I've ever had the pleasure of using! Pod up on thursdai.news/sep-24 and here's a supercut for the unpatient ones
1
2
20
2,720
The Muse models team is telling a really compelling story here at Meta Connect about AI research and product development. 1) The “foundation” of model development is a set of specific product goals. Model releases are paired with product releases. “The model and the product harness must evolve together.” 2) Natively multi-modal, and fundamentally focused on tool use and long context, and on subagent orchestration. This feels like a strategy benefit of restarting a big model project from scratch just a year ago! No legacy approaches to unlearn.
6
11
104
8,382
“Earning the right to act on your behalf” is such a great product North Star.
1
6
729
This is also a *fantastic* slide. “The model and the product weren’t always ready at the same time … we’d been working on these pieces for a long time.”
1
12
591
So ... I went to Meta Connect today expecting to see some cool LLM and AI application demos. And I did see some really good (live) demos. But I also saw the first complete, coherent picture of what an AI-native end-user *platform* looks like. It's been clear for the past two years that generative AI is the next important technology. But the shape of the platform that packages that technology for everybody to use, every day, hasn't quite come into focus. We haven't seen the web browser of the AI era, yet. Today, Meta showed/argued the following: - The Muse AI agent app is "personal super-intelligence." - It's free to use, because eventually Meta will take a tiny piece of every shopping transaction. And our agents are going to buy a lot of stuff for us. - Glasses are the new hardware device form factor. - Voice is the primary interface. - Spatial computing (augmented reality) is here. - The Muse agent is deeply integrated into a new spatial OS. This is "computer use" reimagined from first principles. - Developers can build "connectors" to the Muse agent ecosystem today. And I suspect we'll learn about more developer platform stuff at day 2 of Meta Connect, tomorrow. I attended Meta Connect because there's clearly lots of interesting stuff happening at Meta. The Muse Glimmer and Spark LLMs are quite good. The level of polish and attention to detail in the Muse agent app is something that all of us interested in product design are talking about and dissecting. And in my little corner of the AI world, we were super impressed by the new Muse Spark Transcribe ASR model. But I did not expect to see all these pieces, plus two (or maybe three?) new categories of glasses, stitched together into such a coherent, compelling vision. The glasses: audio-only glasses designed for everyday wear and realtime voice interaction, a huge range of design options, updated camera glasses, and a new version of glasses with built-in display. Plus an entirely new augmented reality glasses form factor that weighs 100g, has hand tracking, and will ship for ~$1,299 next year. The live demo made these glasses look like what I think all of us who are interested in spatial computing were hoping version 3 or 4 of the Apple Vision Pro would evolve into. The tag line for these new glasses is that they are a personal cinema, a computer, and a game console. I actually think this under-sells them. But that's true of how every first-generation transformative product is described. You have to give people analogies to the things they do today, even though once the full ecosystem around a new device gets built out, we don't think of it as a combination of today's x + y + z. The iPhone isn't a phone + camera + mp3 player, even though that's how the original keynote described it. Meta already has content partnerships lined up, of course. (Including sports. There's no such thing as a new platform that doesn't let you watch football, at least in America!) Meta also teased a little keychain-sized "Muse Charm" device, with a microphone and a tiny OLED screen. If you're not wearing your glasses, but you want to talk to your Muse agent, you can pull the charm out of your pocket and interact via voice and by touching the watch-sized screen. Mark Zuckerberg opened the keynote talking about how Meta's north star is believing in people, and building things that empower people. He wore a t-shirt that said "Building is my love language." I've been to a lot of AI keynotes in the past two years. A few of them (Jensen Huang's) have been very, very good. This one was amazing. Not because it had high production values, or wow moments, or showed things that feel like they're an early view of a not-yet-evenly-distributed future. It did have all those things. But mainly, I came away thinking that if Meta is serious about executing on this vision, we'll look back on this keynote as the first time all the pieces of a consumer AI platform came together in one place.
39
37
449
36,579
The official thread:
Here's everything I announced at Meta Connect today 👇
7
1,397
NVIDIA released a fast speaker-labeling model today: Nemotron 3 Diarization. Like all of the Nemotron family from @NVIDIA_AI, this is an open weights model, so you can grab the weights from @huggingface and use the model in the cloud or on your local device. Nemotron 3 Diarization is a useful building block for voice agents and realtime AI applications that need to handle speech from multiple people at the same time. It handles overlapping speech well, can identify up to eight speakers, and pairs with any transcription model you're using in your multi-modal pipelines. @jonptaylor built a demo (code below) and recorded a technical deep-dive showing how to use this diarization model together with a speech-to-text model. Jon's code also uses Jev for super-accurate confirmation of critical voice input in a noisy/multi-speaker environment.
13
14
184
12,247
NVIDIA Nemotron 3 Diarization model card and weights: huggingface.co/nvidia/Nemotr… Jon's demo code: github.com/pipecat-ai/nemo3-… Models used in the demo: - NVIDIA Nemotron 3 Diarization - NVIDIA Nemotron 3.5 ASR - Pipecat Smart Turn - Pipecat PhoneLLM - Typesafe Jev - NVIDIA Magpie TTS
2
2
15
816
The NVIDIA launch post here on X, with links to model benchmarks:
When several people talk at once, a transcript can get messy fast. Our new Nemotron 3 Diarization model tracks who spoke when, even when voices overlap. It handles up to eight speakers, has 100M parameters, and is now available on @huggingface 🤗
2
678
It’s great to see these speech-to-text benchmarks from the Cekura team. Good benchmarks take a *lot* of time to create and maintain. Putting that work in is a valuable community service. Public benchmarks also push all the teams building models and APIs to improve. It’s hard to argue with careful, quantitative benchmark results! Four companies in this benchmark boast a P95 TTFS under 300ms. Cartesia, Deepgram, Gradium, Speechmatics. A year ago nobody had P95 numbers anywhere near that low. We can trace a big jump in TTFS improvement directly to conversations we had about this metric, both in public and in private, when we first published reproducible (full code) realtime transcription benchmarks last year. This surprised me a little bit, because I thought I’d been clearly explaining the importance of “time to final segment” for a while. But code and tables of numbers comparing competitors turns out to be a lot more motivating than just some guy posting words on this 🦅 app. :)
I'm excited to announce that we've launched @cekuraAi's speech-to-text benchmarks: Converse STT, comparing 15 models on transcription accuracy and speed. We tested models from @AssemblyAI, @Google, @OpenAI, @DeepgramAI, @cartesia, @Speechmatics, @inworld, Reson8 and others across public and private datasets. We measured 3 things: 1. Word error rate: how many words the model gets wrong, misses or adds. 2. Time to first text: how quickly it starts returning a transcript. 3. Final-text delay: how long after someone stops speaking the transcript is finalized. For accuracy, we used 1,000 public clips from @pipecat_ai's Pipecat Fleurs dataset and an unseen dataset annotated by our partners at @OcularHQ specifically for this benchmark. And having both datasets gave us a more useful comparison. We've published the results and methodology so you can explore the comparisons yourself: benchmarks.cekura.ai/stt
3
5
36
4,357
Full house at the @general_compute + @SambaNovaAI + @pipecat_ai voice agents hackathon at @AGIHouseSF. Thank you also to the @GradiumAI team for supporting with speech-to-text and text-to-speech APIs.
8
7
40
2,544
Team Pipecat-trainer - drop-in on policy model training for your voice agent
1
215
Team Sandy - elder companions
181