TechnicalAEO Guide

Speaker Diarisation and Audio-Event Annotation for Voice AI

Automatic speech recognition tells you what was said. Speaker diarisation tells you who said it. The annotation discipline for each is different — and voice AI projects that conflate them typically produce training data where the diarisation labels are systematically wrong in ways that ASR transcription annotations will not catch.

3 September 202613 min read

Quick answer

Speaker diarisation annotation is the process of labelling an audio recording with speaker turns — marking segment boundaries at the sub-second level, assigning a unique speaker ID to each turn, and flagging overlapping speech and back-channels. Audio event annotation is the related task of marking onset and offset timestamps for non-speech sounds (alarm tones, background noise classes, music boundaries) with event-type labels. Both produce training data for voice AI systems that need to understand not just what was said but who said it, when, and in what acoustic context. Production diarisation annotation requires annotators skilled in perceptual boundary detection, speaker-ID consistency across multi-hour recordings, and structured handling of overlapping speech — tasks that crowdsourced annotators without explicit training handle poorly, with inter-annotator agreement on overlap segments typically 20–25% lower than for clean turn-taking.

Two Distinct Tasks: Diarisation vs Audio Event Detection

Voice AI annotation encompasses two related but distinct labelling tasks that are often conflated in project scoping — to the detriment of both.

Speaker diarisation annotation answers the question "who spoke when?" It produces a timeline of speaker turns — typically in RTTM (Rich Transcription Time Marked) format — where each segment is labelled with a speaker ID, an onset timestamp, and a duration. The speaker IDs are relative to the recording (Speaker A, Speaker B) and must be consistent across the full file. For multi-session diarisation, IDs may need to be consistent across multiple recordings of the same speakers.

Audio event annotation answers the question "what sounds were present and when?" It produces onset/offset timestamps for non-speech acoustic events — ringing phones, background music, hold tones, machinery noise, applause — each labelled with an event class from a predefined taxonomy. Audio event annotation is used to train acoustic event detection (AED) models for smart speakers, contact-centre telephony analytics, environmental monitoring, and industrial fault detection.

The two tasks share annotation infrastructure (timeline-based tools, timestamp precision requirements) but require different annotator skills and produce different output schemas. Our audio annotation services cover both — and the first step on any voice AI annotation project is determining which task or combination of tasks the model actually needs.

What Makes Diarisation Annotation Hard

Speaker diarisation is one of the perceptually harder annotation tasks in audio. Unlike transcription, where the target is the words, diarisation requires annotators to make boundary decisions at the millisecond level — where exactly does one speaker's turn end and another's begin? — while simultaneously tracking speaker identity across minutes or hours of audio.

Three factors drive most annotation errors:

A 2023 study published in the IEEE/ACM Transactions on Audio, Speech, and Language Processing compared diarisation annotation from three annotator types — expert linguists, trained non-linguist annotators, and crowdsourced annotators — on 120 hours of telephone conversation data. Expert linguists achieved median DER of 4.8% on clean audio. Trained non-linguists achieved 7.3%. Crowdsourced annotators achieved 19.4% — more than four times higher error. The primary driver of the crowdsourced gap was speaker ID inconsistency on long recordings and failure to mark overlapping speech.

Need speaker diarisation or audio event annotation?

Our trained audio annotators deliver RTTM-format diarisation data, timestamped event labels, and combined transcription + diarisation datasets with full QA. Contact-centre, meeting, broadcast, and multilingual recordings covered.

See our audio annotation services

Case Study: Contact-Centre Diarisation for Australian Banking

A major Australian financial services group engaged AI Taggers to build diarisation training data for their contact-centre analytics platform. The existing diarisation system — a pre-trained model without domain fine-tuning — achieved 17.3% DER on their recorded call corpus. The primary failure modes were speaker confusion on accented Australian English and incorrect turn attribution on customer calls with background hold music.

The annotation scope was 480 hours of telephony recordings across four call types: inbound retail banking, inbound home loan, outbound collections, and internal escalation calls. Each recording was annotated by two trained annotators in ELAN, with RTTM export, and adjudicated by a third annotator on segments with boundary disagreement greater than 300ms. Annotators received a four-hour briefing on the bank's specific call structure — including the pattern of agent hold notifications ("I'll just place you on hold") that immediately precede hold music — before annotation began.

Results after fine-tuning on the annotated corpus:

The 480-hour annotation project took nine weeks with a team of six annotators and two senior adjudicators. The improvement in downstream sentiment accuracy — from 71.3% to 88.9% — was the metric the business had been targeting for two years with model architecture changes alone.

The Annotation Pipeline: From Raw Audio to Validated RTTM

A production diarisation annotation pipeline typically runs in four stages:

  1. Pre-segmentation — a voice activity detection (VAD) model segments the audio into speech and non-speech regions, producing a draft timeline that annotators refine rather than create from scratch. Pre-segmentation reduces annotation time by 30–40% but requires annotators to check and correct VAD errors, particularly at the boundaries of overlapping speech and at low signal-to-noise segments.
  2. Speaker turn labelling — annotators assign speaker IDs to each speech segment, marking turn boundaries and flagging overlapping speech. For recordings with more than three speakers, a speaker roster (listing known speakers with brief audio examples from each) is provided to help annotators maintain ID consistency.
  3. Overlap and back-channel marking — a second annotation pass specifically reviews segments where the waveform indicates simultaneous energy from multiple sources. Overlap segments are marked with a dedicated overlap flag and assigned speaker IDs to all concurrent speakers.
  4. Adjudication and QA — disagreements between annotators are resolved; a sample of 5–10% of the corpus is double-annotated by a senior annotator as a gold check; DER is computed on the gold sample to confirm the annotation meets the agreed IAA threshold before the dataset is released.

Audio Event Detection Annotation: Taxonomies and Timestamps

Audio event detection annotation follows a similar pipeline but uses event-class taxonomies rather than speaker IDs. The most widely used reference taxonomy for general-purpose AED annotation is AudioSet (Google, 2017), which provides a hierarchical ontology of 632 audio event classes. For domain-specific annotation, custom taxonomies are typically narrower — a contact-centre AED system might use 12–20 event classes covering hold music, ringing, background speech, keyboard noise, and silence — and annotators must be briefed on the specific taxonomy before annotation begins.

Unlike speaker diarisation, audio event annotation often involves strong labelling (onset and offset timestamps for each event instance) and weak labelling (clip-level binary presence/absence labels without timestamps). Strong labels are more expensive but train better localisation models. Weak labels are cheaper and sufficient for classification tasks that don't need event localisation.

Polyphonic event annotation — where multiple event classes may be simultaneously present — requires multi-track annotation interfaces and explicit handling of the case where one event class masks another. This is the audio equivalent of overlapping speech in diarisation, and it produces similar inter-annotator agreement challenges when guidelines don't specify the minimum event duration to annotate and how to handle masked events.

Cost and Throughput Benchmarks

TaskAnnotation ratioCost per audio hourIAA target (DER/kappa)
Diarisation only (clean audio)3–5× real time$35–$65DER < 8%
Diarisation (noisy/overlapping)6–10× real time$70–$130DER < 12%
Transcription + diarisation8–14× real time$90–$180DER < 8%, WER < 5%
Audio event detection (strong labels)4–6× real time$40–$80kappa ≥ 0.78

Multilingual annotation adds a language-loading factor: annotators who are not native speakers of the recorded language are slower to make turn-boundary decisions because they cannot rely on prosodic and syntactic cues the same way native speakers can. For Australian English recordings, annotators fluent in Australian English are significantly more accurate on boundary detection than annotators working from phonetic cues alone.

Our audio annotation services include multilingual diarisation for Arabic, Mandarin, Hindi, and 40+ other languages — see also our multilingual speech transcription annotation guide for ASR-focused requirements.

QA and Validation: How to Know Your Diarisation Data Is Good

Quality control for diarisation annotation is different from QA for transcription or image annotation, because the primary quality signal is the IAA DER — the diarisation error rate computed between two independent annotators on the same recording — rather than a human-readable label you can inspect visually.

Standard QA protocols for production diarisation annotation:

For a broader view of annotation quality validation methods, see our guide on how to validate annotation quality before it reaches your model — the gold-set and audit approaches apply directly to audio annotation.

Frequently Asked Questions

What is speaker diarisation annotation?+
Speaker diarisation annotation is the process of manually labelling an audio recording with speaker turns — marking exactly when each speaker starts and stops talking, assigning a unique speaker identity to each segment, and flagging overlapping speech and back-channels. The annotated data trains diarisation models that answer the question 'who spoke when?' independently of transcription. Production diarisation annotation requires annotators to mark segment boundaries at the sub-second level, handle overlapping speech, and maintain speaker ID consistency across long recordings.
What is the difference between diarisation annotation and transcription annotation?+
Transcription annotation labels what was said (the words), while diarisation annotation labels who said it and when (speaker identity and segment boundaries). Both are needed for full voice AI training data, but they are distinct tasks with distinct annotator skill requirements. Transcription annotation requires language proficiency and correct spelling; diarisation annotation requires perceptual precision at the boundary level and consistent speaker ID assignment. Many ASR pipelines annotate them in separate passes to maintain quality on each task.
What is Diarisation Error Rate (DER) and what is a production target?+
Diarisation Error Rate (DER) is the standard metric for diarisation system performance, combining speaker confusion, missed speech, and false alarm components. Production call-centre systems typically target DER below 10%. Research state-of-the-art on clean two-speaker telephone conversations achieves DER of 2–5%. In noisy multi-speaker environments with background noise, production DER of 8–15% is common before annotation-guided fine-tuning.
How long does audio annotation take per hour of audio?+
Annotation time varies significantly by task. Speaker diarisation annotation (segment boundaries + speaker IDs only) runs at approximately 3–5× real time. Combined diarisation and transcription runs at 6–10× for clean audio and 10–15× for noisy audio. Audio event detection annotation runs at 4–6× for dense event recordings. These ratios assume experienced annotators working in professional tools.
What audio file formats and tools are used for diarisation annotation?+
Diarisation annotation is most commonly performed in ELAN, Audacity, or Label Studio. Output formats include RTTM (Rich Transcription Time Marked) for diarisation, TextGrid for Praat-based annotation, and WebVTT for subtitle-format outputs. RTTM is the standard format for NIST-aligned evaluations and most diarisation evaluation frameworks including pyannote-metrics.
What makes overlapping speech annotation challenging?+
Overlapping speech — where two or more speakers talk simultaneously — is the hardest annotation case in diarisation. Most tools display a single-track timeline, making simultaneous segments difficult to represent. Annotators must create separate speaker tracks and mark overlap periods on each. Inter-annotator agreement on overlap onset/offset boundaries is typically 15–25% lower than for non-overlapping segments. Guidelines must specify minimum overlap duration to annotate (typically 200ms or longer) and how to handle back-channels.
Free Sample · 24-48 hours

Get a quote for speaker diarisation or audio event annotation

Tell us your audio type (call-centre, meeting, broadcast), language, speaker count, and volume — we'll design the annotation workflow and QA protocol.

No commitment. NDA available on request. We respond within 24 hours, often the same day for Gulf-region inquiries.

Neel Bennett

AI Annotation Specialist at AI Taggers

Neel has over 8 years of experience in AI training data and machine learning operations. He specializes in helping enterprises build high-quality datasets for computer vision and NLP applications across healthcare, automotive, and retail industries.

Connect on LinkedIn