Parakeet Redux is compressed from NVIDIA's Parakeet TDT 0.6B v3, so we ran it against the original, which QVAC runs in Q8.
Both transcribed the same two hours of audio on the GPU of a MacBook Pro M5 Max.
QVAC took 22.3 seconds, 323x real time, and Redux took 44.2 seconds, 163x. Both times include loading the model.
Just released Parakeet Redux!
A ternary speech-to-text model, built by compressing NVIDIA's Parakeet model from 1.2GB to 178MB.
Runs at 113x realtime on CPU, and beats the base model on the 25-language FLEURS benchmark while staying within 0.3 WER on English.