Speechless (insert vocal isolation joke). Thank you @TIME for recognizing AudioShake’s stem separation technology as one of the Best Inventions of 2023. More on why in the year of AI everywhere, our work stood out: time.com/collection/best-inv…
Fun fact for the audio nerds: live broadcast gives you roughly 10–15 milliseconds to process audio before latency becomes a problem.
That’s one reason source separation has historically lived in post-production. Unlike noise reduction, you can actually separate the speech from the background, BUT it’s been too slow.
We’ve been working on changing that at @AudioShakeAI. Dialogue RT separates speech from background audio in 11 ms end to end.
So out of post and into live.
Kinda wild to watch individual sounds disappear in real time.
My super power is that I can guess within 10 seconds of entering any venue / bar / restaurant whether they will play The Shins at some point during my stay.
Dialogue RT isolates dialogue from a live feed in 11 ms end to end — measured model input to isolated output.
That's not noise suppression. It's two stems: dialogue, and everything else.
On NVIDIA DGX Spark and Blackwell architecture GPUs. audioshake.ai/post/isolating…
Four days at IBC, hundreds of conversations. The question we heard most:
“Can it run on the feed, not the file?”
Answer: yes. Live, less than 11 milliseconds. Another great year in the books.
Today we're releasing @AudioShakeAI Multi-Speaker 2.0: one recording of several people talking goes in, a clean labeled track per person comes out — including the moments they talk over each other.
7/
Every job also returns confidence scores, per 20 ms frame and per file: how sure we are the speech went to the right speaker, and how cleanly the overlapping voices came apart.
Rank a million hours, skip the unusable parts.
Coming out of X hibernation to unveil our new brand vision for @AudioShakeAI , courtesy of my 12-yr-old.
P.S. We have a bet on whether I can get more likes than his Brawl Star YT videos, so help me out here.
"Music is more durable than more strictly cognitive tasks like coding."
Why? The core of music lives in live performance, learning an instrument, and the artist-fan connection. The stuff that's hardest to automate is exactly the stuff we value most.
Great clip from @trapital 👇
When the recording is the only copy that exists, you can't let a tool guess. Suppression smears the words. Generative enhancement invents them. So @AudioShakeAI built Speech Recovery to do neither — it isolates the speech that's actually there, and adds nothing.
There is so much media & archives blocked from streaming & social platforms bc of copyright compliance. Distribution demand has multiplied, and the compliance workflows underneath it have not.
Today we're launching the system that closes that gap. tinyurl.com/2nyduuy9
1. Music detection.
2. Music removal (yes---removing music, including with song lyrics--from a mixed media file)
3. Music rights identification
4. Cue sheet creation
A great, practical example of source separation applied to a real-world problem!