STT MOSTLY BEAT BACKGROUND NOISE
But it still loses when another person starts talking nearby.
@krispHQ just open-sourced its 'Krisp Voice Isolation' Benchmark on @HuggingFace to test exactly that:
how much can voice isolation improve speech-to-text?
> 265 real recordings
> 11 STT setups
> Word error rate: 23.3% → 6.2% with isolation
That’s roughly a 73% reduction.
The gains were even bigger in shared offices and call-center environments.
Phone calls barely moved, which makes sense.
There’s usually much less competing speech to remove.
→ real recordings, not synthetic mixes
→ complete dataset on Hugging Face
→ fully reproducible
Krisp Voice Isolation already powers 200+ voice platforms and has processed 10B+ minutes of voice-agent audio.
Now the benchmark data is open 🤗
(link in 🧵 ↓)
Paid partnership (ad)
5
22
69
9,497
check it out here > partner.krisp.ai/datachaz-x-…
weights on HF > huggingface.co/spaces/Krisp-…
1
3
4
843
Big thanks to @krispHQ for the collab 🙏
If you found this useful, a like or repost helps get it in front of more people 🦾
Follow me here or on 𝕏 → @datachaz for more LLMs, AI agents, and data science.
STT MOSTLY BEAT BACKGROUND NOISE
But it still loses when another person starts talking nearby.
@krispHQ just open-sourced its 'Krisp Voice Isolation' Benchmark on @HuggingFace to test exactly that:
how much can voice isolation improve speech-to-text?
> 265 real recordings
> 11 STT setups
> Word error rate: 23.3% → 6.2% with isolation
That’s roughly a 73% reduction.
The gains were even bigger in shared offices and call-center environments.
Phone calls barely moved, which makes sense.
There’s usually much less competing speech to remove.
→ real recordings, not synthetic mixes
→ complete dataset on Hugging Face
→ fully reproducible
Krisp Voice Isolation already powers 200+ voice platforms and has processed 10B+ minutes of voice-agent audio.
Now the benchmark data is open 🤗
(link in 🧵 ↓)
Sep 24, 2026 · 7:38 PM UTC
1
674

