Skip to main content
Pre-recorded Speaker diarization assigns speaker labels to segments of pre-recorded audio so you can identify who spoke when. In Gladia, enable diarization on the pre-recorded transcription request. The response associates each utterance with a speaker index in order of first appearance. These labels distinguish speakers within the recording and do not establish a person’s identity.

Enabling diarization

Diarization is enabled by sending the diarization parameter in the transcription request:
Pre-recorded

Response

When diarization is enabled, each utterance will contain a speaker field, whose value is an index representing the speaker. Speakers will be assigned indexes by order of appearance (i.e. the 1st speaker will be speaker 0, the 2nd speaker 1, etc).
Pre-recorded

Improving diarization accuracy

Provide speaker-count hints with diarization_config.number_of_speakers, diarization_config.min_speakers and diarization_config.max_speakers. These specify the expected count, lower hint and upper hint respectively. They are hints, not hard constraints; the detected count may differ.

Diarization scope and evaluation

This guide covers speaker diarization for pre-recorded audio only. Speaker diarization labels who spoke when within a single mixed audio track, while transcription accuracy measures word errors in the transcribed text. Channel identification instead relies on separate audio channels, reported in each utterance’s channel field, rather than telling speakers apart within one track (see Multiple channels). For concept definitions, see speaker diarization concepts. Async accuracy comparisons use the model and dataset scope described in the async benchmark methodology. Use the blind API comparison tool and check pricing for current plans. Speaker-count hints do not guarantee a specific detected speaker count.