End-of-turn detection is solved with our SOTA v1 model.
It fuses an audio encoder directly into the LLM backbone, so the model can use both audio cues and semantic understanding together. It's the fastest and most accurate turn detection model we've tested.
We shipped LiveKit Turn Detector v1.
Instead of reading transcripts, it listens to speech directly, combining semantic and acoustic cues into one end-of-turn prediction.
The result: high accuracy, low latency—the best model we tested across 14 languages.
Available on LiveKit Cloud.