Multimodal Inference at the Speed of Light. Try now at narilabs.com

San Francisco, CA
Word timestamps are valuable - but can come at a cost. With clever inference engineering, we keep performance same while adding a 0.6B model to the mix.
Just added word timestamps support to our Qwen3-ASR streaming endpoint - using the 0.6B forced aligner model. Still the same latency and cost! Blog: narilabs.com/blog/word-level…
4
125
Input streaming is a game changer for voice agents. Now live at our inference platform!
heavily requested feature: input streaming for Qwen3-TTS is now GA! stream your LLM outputs into our TTS endpoint so you can output audio before the LLM is finished ✅
4
73
The era of slow and expensive TTS and STT is coming to an end. Try now at narilabs.com 🚀
Our STT and TTS endpoints are entering GA today 🥳 Try now and receive welcome credits for testing! > TTS Fast: $10 / 1M (top TTFA, 50 ms) > TTS Standard: $5 / 1M > STT Fast: $0.12 / hr (top TTFS, 40 ms) > STT Standard: $0.06 / hr
4
209
Below 10 cents an hour to serve a single user? The future of interfaces will be voice. And we’re making it come true!
Duplex speech is the frontier of Speech AI and PersonaPlex 7B from @nvidia is the most adopted OSS duplex model. We achieve 80 concurrent sessions on a single H100 GPU, costing below 10 cents an hour per user to serve. GPT-Live costs 10 cents per 2 minutes. Read more: narilabs.com/blog/personaple…
2
144
We will be launching Diarization endpoints soon. Stay tuned!
Diarization is a key problem in speech. And Pyannote is the king of diarization. Used it for all of Dia's data pipelines. Through various optimizations, we got Pyannote's flagship OSS model to process 100+ hours of audio in under 100 seconds. on a single A100 GPU. Read more: narilabs.com/blog/making-pya…
1
5
242
Crazy numbers. Try now at narilabs.com
Nari Labs' Qwen3-TTS and Qwen3-ASR endpoints lead @covaldev voice AI benchmarks 🎉 Text-to-Speech > #1 Lowest Latency (time-to-final-segment): 44 ms > #2 Accuracy: 3.6% WER, just 0.1% shy of #1 Speech-to-Text > #1 Accuracy: 3.8% WER > #2 Lowest Latency (time-to-first-audio): 63 ms Read more: narilabs.com/blog/nari-labs-…
7
371
Nari Labs retweeted
Today we’re launching: > the world's fastest Speech-to-Text endpoint at 40 ms time-to-final-segment (TTFS) > the world's cheapest streaming Speech-to-Text endpoint at $0.06 / hr Powered by Qwen3-ASR 1.7B on Nari Labs inference engine. 🔠 Top 3 Accuracy on Coval Benchmarks 🚀 3x faster than ElevenLabs Scribe V2 💸 9x cheaper than Gemini Transcribe 3.5 ✅ available FREE for a limited time We believe open-source will win - not just in LLMs, but also in multimodal AI. Nari Labs is here to accelerate that future, starting with speech. Try now at narilabs.com
26
35
413
23,474
Nari Labs retweeted
Today we're launching > the world's fastest TTS endpoint at 50 ms time-to-first-audio (TTFA) > the world's cheapest modern TTS endpoint at $5 / 1M characters Powered by Qwen3-TTS 1.7B on Nari Labs inference engine. 🚀 5x faster than Cartesia 💸 10x cheaper than ElevenLabs 🎙️ expressive voices, outperforming industry average 🎁 available FREE for a limited time We believe open-source will win: not just in LLMs, but also in multimodal AI. Nari Labs is here to accelerate that future, starting with speech. Try now at narilabs.com
46
48
608
43,505