Arabic & MENA

Choosing an Arabic Data Annotation Company: 10 Checks Before You Sign

Most Arabic annotation failures are procurement failures. The right checks happen before the contract, not after the data comes back wrong.

10 October 2026 13 min read

Direct Answer

An Arabic data annotation company is a specialist provider that labels Arabic text, audio, or image data using verified native-speaker annotators segmented by dialect — Khaleeji, Egyptian, Levantine, and MSA. The right vendor proves dialect coverage before the contract, reports inter-annotator agreement by task type, handles PDPL data-residency obligations, and offers a pilot of 200–500 samples before any volume commitment.

Why Generic Annotation Vendors Fail at Arabic

Arabic is spoken by over 400 million people across 22 countries (Ethnologue, 2024), but it is not a single language. It operates as a diglossic system: Modern Standard Arabic (MSA) is used in formal writing, news, and government documents, while regional spoken dialects — Khaleeji, Egyptian, Levantine, Moroccan Darija — differ substantially in vocabulary, grammar, and pragmatics. A native speaker of Egyptian Arabic does not reliably label Emirati colloquial chat transcripts at production-grade accuracy.

Generic annotation platforms address this by recruiting bilingual annotators — people with strong English and formal Arabic literacy. For classification tasks on MSA news text, this is serviceable. For conversational AI, sentiment analysis on social media, or any task requiring dialect judgement, the gap becomes significant.

Published NLP benchmarking research has consistently found that named entity recognition and intent classification models trained on MSA-annotated data show substantially lower F1 scores on Gulf dialect test sets compared to models trained on dialect-native annotated data. The gap is not a modelling artefact — it is a labelling provenance problem.

Our Arabic data labelling service works from a structured annotator pool where each annotator carries a verified dialect tag, has passed a task-specific qualification test, and is calibrated against gold sets for ongoing quality measurement. The 10 checks below are designed to surface whether any vendor you evaluate has that same infrastructure in place.

The 10 Checks Before You Sign

These checks apply whether you are buying NLP annotation, speech transcription, sentiment labelling, or multimodal Arabic annotation. Request written responses to each — verbal assurances rarely survive the handover to the delivery team.

01

Dialect Coverage Map

Ask the vendor to show their annotator pool breakdown by dialect — Khaleeji (broken further into Emirati, Saudi, Qatari where relevant), Egyptian, Levantine, MSA, and Darija. A vendor that claims 'Arabic' without a dialect breakdown cannot guarantee that Gulf-market data is routed to Gulf-native annotators. Minimum viable threshold for a Khaleeji project: at least 30 active annotators with verified Gulf-dialect provenance.

02

Native-Speaker Vetting Process

Ask how annotators are recruited and qualified. Look for: a dialect identification test at onboarding, a task-specific qualification assessment for your annotation type, and a minimum score before live data is assigned. Reject vendors who rely entirely on self-reported native-speaker status — dialect competency tests must be in the process, not optional.

03

Inter-Annotator Agreement (IAA) Reporting

Ask for IAA scores — Cohen's kappa, Fleiss's kappa, or Krippendorff's alpha — by task type from recent comparable projects. Production-grade Arabic NER: κ ≥ 0.80. Sentiment on dialect text: κ ≥ 0.75. Vendors that cannot produce per-task IAA reports have no reproducible quality baseline. Ask specifically for dialect-level IAA, not an aggregate figure.

04

PDPL and Data Residency

If your data contains personal information about Saudi residents, PDPL requires a data processing agreement before transfer, restrictions on cross-border movement, and documented deletion at project end. Ask: where is annotation work performed? Do they have a DPA template for GCC clients? How is data deleted? Similar requirements apply under UAE PDPL, DIFC DPL, and ADGM DPR.

05

RTL Tooling

Ask which annotation platform they use for Arabic text. It must handle right-to-left rendering natively, support the full Arabic Unicode block (including diacritics and lam-alef ligatures), and render entity span labels without direction artefacts. Ask for a screenshot of an Arabic NER task in their tooling before you proceed.

06

Code-Switching Protocol

GCC business communications commonly mix Arabic and English — a customer message might start in Khaleeji and switch to English for product names or technical terms. Ask how the vendor handles code-switching: do they label spans by language, route tasks to annotators fluent in both, and have documented guidelines for mixed-language intent classification?

07

Adjudication Process

When two native-speaker annotators disagree, what happens? A proper adjudication layer — a senior annotator or domain expert reviews the disagreement and makes a binding decision — is what separates IAA as a QA gate from IAA as a vanity metric. Ask to see the escalation workflow, including who the adjudicators are and how they are qualified.

08

Gold Set QA

Does the vendor maintain task-specific gold sets — held-out samples with known-correct labels — and report annotator accuracy against them? Ask for their gold set construction methodology: who created the gold labels, and were they produced by senior native speakers or imported from a public dataset, which may carry its own dialect biases?

09

Pilot Flexibility

Will they annotate 200–500 samples before you commit to volume? Any vendor requiring a minimum volume commitment before a pilot cannot demonstrate quality at small scale. A structured pilot should include: the labelled output, IAA scores, annotator profile (dialect breakdown), delivery in your target schema, and a written retrospective on guideline edge cases.

10

Security Controls

For sensitive data — financial records, healthcare data, government documents — ask about ISO 27001 certification, annotator NDA coverage, managed-device requirements (no personal laptops), no-public-tool policies (data must not be passed to ChatGPT, Google Translate, or similar), and audit trail capabilities for regulatory provenance.

Case Study: UAE Digital Services Platform

Project Snapshot

50,000
Records
MSA + Emirati
Dialects
Intent + NER
Task
6 weeks
Timeline

A government-aligned digital services platform in the UAE needed to annotate 50,000 citizen-service chat transcripts in a mix of MSA and Emirati dialect for a virtual assistant rollout. Their first vendor — a generalist platform using bilingual but non-native annotators — returned data with an inter-annotator agreement of 68% on intent labels and a 22% error rate on dialect-specific named entities. Neither figure was flagged by the vendor. The problems only surfaced when the downstream model underperformed significantly on held-out UAE test data.

After switching to a native-speaker annotation pipeline with per-dialect routing and a double-blind adjudication layer, IAA improved to 91% on intent classification and 88% on NER. The dialect-specific entity error rate dropped from 22% to 4%. The downstream virtual assistant achieved an 18-point improvement in intent accuracy on held-out Emirati citizen queries.

The post-mortem finding was simple: the original vendor had no Emirati-native annotators at all. The work had been routed to MSA-trained annotators who consistently miscategorised Emirati-specific lexical patterns. None of the 10 checks above had been applied at procurement.

Start With a Free Arabic Annotation Pilot

Send us 200–500 Arabic records across your target dialects. We will annotate them free and return the output with full IAA reporting within 48 hours — so you can verify quality before any volume commitment.

Request a Free Arabic Annotation Pilot

PDPL, Data Residency, and What Most Vendors Miss

Saudi Arabia's Personal Data Protection Law (PDPL) came into full effect in September 2023 and is enforced by SDAIA (the Saudi Data and Artificial Intelligence Authority). The law applies to any organisation — domestic or foreign — that processes personal data about Saudi residents. For annotation vendors, this creates specific obligations that many offshore platforms have not yet structured their Middle East operations to address:

  • A data processing agreement (DPA) must be signed before any personal data is transferred to the vendor
  • Cross-border data transfers require either explicit data-subject consent or a SDAIA-approved transfer mechanism — or annotation work must be performed within the Kingdom
  • Data must be deleted or irreversibly anonymised at the end of the engagement, with documented evidence of deletion
  • Breach notification to SDAIA is required within 72 hours of discovery

Parallel requirements apply for UAE data: the Federal UAE Personal Data Protection Law (enacted 2021) applies nationally, with the Dubai International Financial Centre Data Protection Law (2020) and Abu Dhabi Global Market Data Protection Regulations applying in their respective free zones.

Many global annotation platforms handle GDPR compliance competently but have not structured their operations for PDPL. When evaluating vendors, ask specifically: where is annotation work physically performed? Do they have signed DPAs with GCC clients on record? What is their deletion process at project end? For healthcare, financial, or government data, request evidence of compliance process — not just a security checkbox on a vendor portal.

What a Useful Pilot Includes

A pilot is your quality audit at manageable scale. For Arabic annotation, a useful pilot delivers more than labelled data — it reveals the quality infrastructure. A well-structured Arabic annotation pilot should include:

200–500 labelled records from your actual data (not generic demo samples)
Annotator profile: number of annotators, dialect breakdown, qualification scores
IAA scores by task type and dialect (not aggregate only)
Delivery in your target schema (JSON, JSONL, CSV with your field names)
A written retrospective on guideline edge cases discovered during the pilot
A sample of adjudicated disagreements showing how edge cases were resolved

For audio tasks, also request: timestamp accuracy on a 30-second sample clip, a speaker-diarisation example for a two-speaker conversation, and the transcription output in both buckwalter and Unicode Arabic formats if your pipeline requires both. See our guide to end-to-end Arabic data labelling pipelines for how production annotation workflows are structured from task brief to delivery.

Cost Drivers and Realistic Turnaround

Arabic annotation pricing varies by task complexity, dialect requirement, and quality tier. The major cost drivers are dialect routing (Khaleeji and Darija specialists command a premium over MSA annotators), adjudication layers (adds approximately 25–35% to base cost), domain expertise (clinical and legal Arabic requires subject-matter-qualified annotators), and volume (per-unit rates drop substantially above 50,000 records). For specific pricing, see our annotation pricing page.

Turnaround for a 10,000-record Arabic text annotation project is typically 5–7 business days at standard quality. Dialect-specific work requiring adjudication and gold-set calibration takes 8–12 days. Audio transcription at 10,000 minutes of Gulf-dialect speech typically takes 10–15 days with native-speaker QA.

For further reading on Arabic annotation tooling and software evaluation, see our guide to Arabic text annotation software. For Gulf-dialect sentiment specifically, see our deep dive on Khaleeji Arabic sentiment annotation. If you are building an LLM on Arabic data, our Arabic text annotation service covers SFT and RLHF data preparation workflows.

Frequently Asked Questions

What is an Arabic data annotation company?

An Arabic data annotation company is a specialist service provider that labels Arabic text, audio, image, or video data for AI training. The defining characteristic is a curated pool of native-speaker annotators segmented by dialect — not just bilingual generalists. Quality providers cover at minimum Khaleeji, Egyptian, MSA, and Levantine Arabic, and report inter-annotator agreement by dialect and task type.

What dialects should an Arabic annotation vendor cover?

At minimum: MSA, Gulf/Khaleeji (UAE, Saudi Arabia, Qatar, Kuwait, Bahrain), Egyptian Arabic, and Levantine. For MENA-market products, Moroccan Darija and Iraqi Arabic are increasingly important. Any vendor that only claims 'Arabic coverage' without dialect detail should be asked to specify annotator pool composition before you proceed.

Does my Arabic annotation vendor need to be PDPL-compliant?

Yes, if the data contains personal information about Saudi residents. PDPL (enforced by SDAIA since September 2023) requires a data processing agreement before transfer, restrictions on cross-border movement, and documented deletion at project end. UAE data is governed by the Federal UAE PDPL and, in the DIFC and ADGM free zones, by additional local regulations.

What is a good IAA score for Arabic NLP annotation?

Cohen's kappa above 0.80 is production-grade for Arabic NER and intent classification. Sentiment on dialect text typically achieves 0.75–0.82. Always ask for IAA broken down by dialect, not as an aggregate — MSA IAA tends to be higher and can mask problems in the dialect-specific portions of your data.

How long does an Arabic annotation pilot take?

For text tasks (NER, intent, sentiment), a 200–500-record pilot with IAA reporting takes 48 to 72 hours. Audio transcription pilots take 3 to 5 business days. Allow additional time if dialect identification is required before routing, or if the vendor needs to draft annotation guidelines from scratch.

Free Sample · 24-48 hours

Get a Free Arabic Annotation Pilot

Send us 200–500 records across your target dialects. We'll annotate them free and return the results with full IAA scores within 48 hours.

This form is for companies with annotation projects. Looking for annotation work? Apply on our careers page. Job enquiries sent here don't get a reply.

No commitment. NDA available on request. We respond within 24 hours, often the same day for Gulf-region inquiries.

Neel Bennett

AI Annotation Specialist at AI Taggers

Neel has over 8 years of experience in AI training data and machine learning operations. He specializes in helping enterprises build high-quality datasets for computer vision and NLP applications across healthcare, automotive, and retail industries.

Connect on LinkedIn