Diarization, Recognition, Identification, Verification.
Four terms, four different problems. Confusing them produces architecture mistakes that surface late, usually in production.
What each one actually does, and where it sits in a Voice AI pipeline ↓
1
1
8
1,158
Speaker Diarization answers "who spoke when."
It partitions audio into segments and groups them by voice, using anonymous labels: SPEAKER_00, SPEAKER_01.
No enrollment, no prior knowledge of the speakers. Pure acoustics: pitch, timbre, rhythm.
1
1
170
Voice print is a compact fingerprint, a fixed-size embedding derived from someone's voice.
Creating one involves no comparison at all. It is just a representation.
The comparison is what turns it into a capability, and there are two ways to run it.
1
50
Speaker Verification → "Is this really Alice?" Binary answer against a single enrolled profile.
Worth stating plainly: voiceprint verification is a soft authentication signal or anomaly detector. Not a biometric gate like face unlock.
1
29
Speaker identification → "Which of our enrolled speakers is this?" No claimed identity needed upfront. Returns an identity label with a confidence score.
1
23
The detail most teams miss:
Identification always performs diarization at the same time. One request returns both the segment timeline and the resolved identities.
You do not need a separate diarization call. Verification, by contrast, returns no timeline.
1
26

