Clean audio is easy. Real conversations aren’t.
People switch languages halfway through a sentence. They change their minds, speak over traffic noise, use bad microphones — and still expect the model to keep up.
So we want to test the messy stuff.
Send us a short, non-sensitive audio sample or tell us about a difficult voice scenario. We’ll test it with Hojo-ASR-Multi-V1 and share what works, what doesn’t, and where the model still needs improvement.
What should we test first?
Model:
huggingface.co/HojoAI/Hojo-A…