Cheng-I Jeff Lai retweeted
NVIDIA Nemotron 3 Diarization is now available on ModelScope—adding live speaker attribution to existing ASR workflows without replacing the transcription model. 🎙️ 🤖 modelscope.ai/models/nv-comm… 👥 Processes streaming audio and returns speaker labels and timestamps for up to eight speaker slots in a single conversation. ⚡ Its end-to-end streaming architecture avoids separately combining voice activity detection, speaker embeddings, clustering, and post-processing. 🧠 The 99.2M-parameter model uses a 31-layer Transformer encoder with RoPE and builds on NVIDIA’s Streaming Sortformer architecture. 🔌 Pair it with Nemotron ASR, Parakeet, Canary, Whisper, or another ASR system to create speaker-attributed transcripts. 🏢 Designed for meetings, contact centers, clinical conversations, live captioning, media analysis, and multi-party voice agents. 🖥️ Supports NVIDIA Ampere, Hopper, and Blackwell GPUs, with inference through NeMo Speech C++.
Made with AI
3
10
71
4,902
Cheng-I Jeff Lai retweeted
Big update: @WaveFormsAI was acquired by @Meta. Today at #MetaConnect, you’ll see a little of what we’ve all been building. ✨🤖
21
27
343
50,133
Even in chaotic moments, people can be kind
1
14
3,612
Data is more important than ever
We've raised $100M from Kleiner Perkins, Index Ventures, Lightspeed, and NVIDIA. Today we're introducing Sonic-3 - the state-of-the-art model for realtime conversation. What makes Sonic-3 great: - Breakthrough naturalness - laughter and full emotional range - Lightning fast -
5
5,345
Speech‑to‑speech generation eliminates the need for separate (TTS) text‑normalization and prosody modeling, both of which are notoriously hard.
1
1
14
2,165
Free will, so often overlooked, becomes priceless only when it slips away.
2
1,435