Make the most of conversational speech with state of the art voice AI

France
🎙️ We’re launching the official pyannoteAI Community Discord A place for developers, researchers, and Voice AI enthusiasts to: ✔ Get support & share knowledge ✔ Showcase projects & learn from each other ✔ Discuss trends shaping the Voice AI ecosystem discord.gg/vux8UH9QmV
4
2,775
Overlapping speakers break your transcripts. The latency budget is gone once you chain models. Costs climb at scale. If that's familiar, come talk about it in Paris on Oct 14. pyannoteAI x Gladia x Modal luma.com/paris-voice-ai
46
One week until our live session on Precision-3! October 8, 5 PM CEST. Hervé Bredin and Jyoti Bisht will walk through the benchmark results against Precision-2, the new operating modes, the input parameters, sensitivity controls, and probability outputs.
1
659
If you're running diarization in production and want to know exactly what changed before you upgrade, this is the session.
1
65
If you no longer believe in the benchmarks published by various Voice AI model providers, this webinar is for you, as we’ll share Precision-3’s results without any hype or false claims. Register here: app.livestorm.co/pyannote-ai…
75
Most diarization benchmarks aren't comparable across providers. Precision-3 uses the same 15-dataset methodology as Precision-2: 10.4% lower error, mostly from fixing wrong-speaker attribution. Full results, including where we don't win. 📊 Explore it: pyannote.ai/blog/precision-3…
2
5
794
Voice AI Meetup, Paris, Oct 14. pyannoteAI x Gladia x Modal on what it takes to run diarization, transcription, and infra in production - not a demo. Panel + Q&A, then drinks. luma.com/paris-voice-ai
2
3
219
- VAD_sensitivity goes from -0.8 on restaurant audio to +2.0 on audiobooks, so no single value wins everywhere. - Crosstalk_sensitivity has a cliff: stay between -1 and 0 for general diarization; pushing it higher makes the model flood the output with overlap that isn't there.
2
2
246
We include some settings that you can copy into your pipeline for transcription, voice cloning, dataset prep, noisy multi-speaker audio, and STT reconciliation.
1
67
This post also reflects how we want to work at pyannoteAI. Our researchers are committed to writing accessible content that shares their expertise and shows how to use our diarization models in production, beyond benchmark scores. Full post: pyannote.ai/blog/tune-vad-cr…
51
Reading about Precision-3's new parameters is one thing. Knowing when to use them is another. Our DevRel team walks through every feature on real audio, from tuning vadSensitivity to reading the new probability scores. ➡️ piped.video/gmhNJixkQw0
2
1
14
1,648
What counts as good diarization changes by team. Some need every overlapping stretch flagged, even if that means a few false alarms. Others need short pauses broken out precisely. Some would rather work with fewer, more dependable speaker segments than a complete but noisier set.
2
2
183
We'll cover: - What changed in Precision-3 and why we built it - The benchmark methodology behind the accuracy and speaker assignment claims - The sensitivity controls and per-frame probabilities and the speed, balance and accuracy modes. - How to migrate from Precision-2
1
40
pyannoteAI retweeted
2 weeks in and I’ve already shipped docs and videos for our latest launch - precision-3 at @pyannoteAI This is the acceleration and speed you get to work with at startups and I’m all here for it. Watch the latest tutorial to see what’s new : piped.video/gmhNJixkQw0?si=pAmJ… Nerd out in comment sections about voice models!
3
33
1,485
pyannoteAI retweeted
Many of our customers run workloads with speaker attribution, with the strictest demands on both quality and speed. We partnered with the team at @pyannoteAI to deliver both: some of the highest-quality diarization models on the market, with 3.2x higher throughput and 9.6x lower latency.
3
3
18
1,812
Precision-3 is live. Diarization error down 10.4% vs Precision-2 (14.35 vs 16.02 DER across 15 benchmark datasets), with most of that gain coming from better speaker attribution: the model gets the wrong speaker 18.7% less often.
3
3
9
278
It also comes with the controls to act on diarization results: two input parameters (vadSensitivity, crosstalkSensitivity) and three output scores (speakerProbability, speechProbability, crosstalkProbability). To tune detection behavior and read confidence per frame.
1
101
If you're running diarization in production, let’s give Precision-3 a try. Here is everything you need to learn about this new model ➡️: pyannote.ai/blog/precision-3
3
78