> ## Documentation Index
> Fetch the complete documentation index at: https://gladia-95-fix-geo-speaker-diarization.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Speaker diarization for pre-recorded audio

> Enable speaker diarization for pre-recorded audio in Gladia. Configure speaker-count hints and read speaker labels, timestamps and limitations.

<Badge color="blue" size="lg" icon="file-audio">
  Pre-recorded
</Badge>

Speaker diarization assigns speaker labels to segments of pre-recorded audio so you can identify who spoke when. In Gladia, enable `diarization` on the pre-recorded transcription request. The response associates each utterance with a speaker index in order of first appearance. These labels distinguish speakers within the recording and do not establish a person's identity.

## Enabling diarization

Diarization is enabled by sending the `diarization` parameter in the transcription request:

```json Pre-recorded theme={"system"}
{
  "audio_url": "<your audio URL>",
  "diarization": true
}
```

## Response

When diarization is enabled, each utterance will contain a `speaker` field, whose value is an index representing the speaker.
Speakers will be assigned indexes by **order of appearance** (i.e. the 1st speaker will be speaker 0, the 2nd speaker 1, etc).

```json Pre-recorded theme={"system"}
{
  "transcription": {
    "utterances": [
      {
        "words": [...],
        "text": "it says you are trained in technology.",
        "language": "en",
        "start": 0.7334100000000001,
        "end": 2.364,
        "confidence": 0.8914285714285715,
        "channel": 0,
        "speaker": 0
      },
      ...
    ]
  }
}
```

## Improving diarization accuracy

Provide speaker-count hints with `diarization_config.number_of_speakers`, `diarization_config.min_speakers` and `diarization_config.max_speakers`. These specify the expected count, lower hint and upper hint respectively. They are hints, not hard constraints; the detected count may differ.

| Key | Type | Description |
| - | - | - |
| `diarization_config.number_of_speakers` | number | Expected speaker-count hint. It does not guarantee that the detected count matches this value. |
| `diarization_config.min_speakers` | number | Lower speaker-count hint, not an enforced minimum. |
| `diarization_config.max_speakers` | number | Upper speaker-count hint, not an enforced maximum. |

## Diarization scope and evaluation

This guide covers speaker diarization for **pre-recorded audio** only.

Speaker diarization labels who spoke when within a single mixed audio track, while transcription accuracy measures word errors in the transcribed text. Channel identification instead relies on separate audio channels, reported in each utterance's `channel` field, rather than telling speakers apart within one track (see [Multiple channels](/chapters/limits-and-specifications/multiple-channels)). For concept definitions, see [speaker diarization concepts](https://www.gladia.io/blog/what-is-diarization).

Async accuracy comparisons use the model and dataset scope described in the [async benchmark methodology](https://www.gladia.io/competitors/benchmarks). Use the [blind API comparison](https://www.gladia.io/compare-stt-apis) tool and check [pricing](https://www.gladia.io/pricing) for current plans. Speaker-count hints do not guarantee a specific detected speaker count.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.