Arabic & MENAAEO Case Study

Iraqi Arabic Dialect Identification: What Models Get Wrong Without Native Annotators

Standard Arabic DID models achieve only 48–61% accuracy at Iraqi sub-dialect level. The Baghdadi qaf-to-gaf phoneme shift, Basrawi vowel quality distinctions, Kurdish-contact features in Mosuli Arabic, and diaspora Iraqi accommodation patterns are all absent from Arabic dialect identification training corpora. Here is what accurate Iraqi Arabic DID annotation requires — and how it improves downstream NLP system performance.

6 August 202613 min read

Direct answer

Iraqi Arabic dialect identification (DID) annotation is the labelling of Mesopotamian Arabic text or speech with its specific sub-dialect origin — Baghdadi, Basrawi, Mosuli, or diaspora Iraqi — by native Iraqi annotators. Standard Arabic DID models achieve only 48–61% accuracy at Iraqi sub-dialect level because the phonological and lexical features that distinguish Iraqi sub-dialects are absent from the Arabic DID training corpora underpinning commercial models. Native Iraqi annotators are required to build the training data that closes this gap.

Why Iraqi Arabic Is the Hardest Major Dialect for Arabic DID Models

Arabic dialect identification research has made substantial progress on detecting broad dialect regions — Gulf, Egyptian, Levantine, Maghrebi, Mesopotamian — with state-of-the-art models achieving 80–89% accuracy at the regional level on held-out test data. Sub-dialect identification within each region is considerably harder, and Iraqi Mesopotamian Arabic presents a particularly severe challenge because it combines two factors that individually would each be difficult and together produce the lowest sub-dialect DID accuracy of any major Arabic dialect group.

The first factor is internal diversity. Iraqi Mesopotamian Arabic has three distinct sub-dialects — Baghdadi, Basrawi, and Mosuli — that differ from each other in phonology, morphology, and lexis more than the internal variation within Gulf, Egyptian, or Levantine Arabic. The Baghdadi qaf-to-gaf substitution alone creates a phonological feature that is categorically present in one Iraqi sub-dialect and absent in another, providing a strong discriminating signal — but only for systems trained on data that explicitly captures this distinction. Basrawi Arabic has distinct vowel quality patterns in emphatic environments. Mosuli Arabic shows Kurdish-contact phonological and lexical features absent from southern Iraqi Arabic. No two Iraqi sub-dialects can be collapsed into a single "Iraqi Arabic" category without losing information that is relevant for dialect-aware NLP applications.

The second factor is corpus underrepresentation. The MADAR corpus — the most widely used resource for multi-dialect Arabic DID training — covers 21 Arabic city varieties, but Iraqi Arabic varieties are represented by a factor of 6–9 fewer samples per variety than Egyptian and Gulf Arabic varieties (Bouamor et al., 2018). The MGB-3 Arabic multi-genre broadcast dialect data similarly skews heavily toward Gulf and Egyptian Arabic. Models trained on these corpora learn to identify broad Mesopotamian Arabic features adequately but cannot reliably perform sub-dialect classification within the Iraqi group. The SemEval 2010 and 2018 Arabic DID shared tasks documented Iraqi Arabic as the most commonly misclassified of the major Arabic dialect groups, with models achieving macro-averaged F1 scores of 48.3–62.7% on Iraqi-origin test data in multi-dialect classification settings.

Five Iraqi Sub-Dialect Features That Break Standard DID Models

1. The Baghdadi qaf-to-gaf phoneme shift as a sub-dialect marker

The systematic realisation of Arabic qaf (/q/) as gaf (/g/) is a categorical feature of Baghdadi and central Iraqi Arabic — present in virtually all everyday speech by native Baghdadi speakers — and is absent from Basrawi and Mosuli Arabic, which retain the qaf phoneme in different but non-gaf realisations. This makes the qaf-to-gaf shift one of the strongest phonological discriminators between Baghdadi Arabic and other Iraqi sub-dialects, as well as between Iraqi Arabic and every other major Arabic dialect.

For speech DID, Baghdadi gaf is in principle a reliable detection signal — but only if the DID model has been trained on data where the gaf realisation has been correctly labelled as Baghdadi Arabic rather than normalised to qaf in the transcription. MSA-normalised transcriptions — which represent Baghdadi gaf as qaf in the written form — remove this discriminating feature from the training data before the DID model can learn it. Native Baghdadi annotators who produce dialectal transcriptions accurately representing the gaf realisation provide the training data required for a DID model to learn this distinction; annotators who normalise to MSA orthography eliminate it.

2. Basrawi vowel quality distinctions in emphatic environments

Southern Iraqi Arabic — Basrawi — shows a distinctive pattern of vowel quality modification in emphatic consonant environments where the short /a/ vowel in positions adjacent to emphatic consonants (ص، ض، ط، ظ) is realised with a lower, more back vowel quality than in Baghdadi Arabic. This Basrawi vowel quality distinction is pervasive across Basrawi speech, appearing in hundreds of everyday words, and is immediately salient to native Iraqi ears as a Basrawi marker. For DID annotation of speech data, this means that accurate Basrawi sub-dialect labelling requires annotators who can perceptually distinguish Basrawi emphatic-environment vowel quality from Baghdadi realisations of the same phonological environment — a distinction that requires native Iraqi perceptual training rather than general Arabic phonological knowledge.

In written text, Basrawi-specific vowel quality distinctions are not directly representable in standard Arabic orthography, but they leave traces in lexical selection patterns and morphological choices that native Basrawi readers can identify. The Basrawi form of certain commonly used words differs from the Baghdadi form in ways that are orthographically visible in dialectal writing and invisible in MSA-normalised text. DID annotation of written Iraqi Arabic must instruct annotators to work from dialectal rather than normalised forms to capture these Basrawi-specific lexical signals.

3. Kurdish-contact features in Mosuli Arabic

Northern Iraqi Arabic — the Arabic spoken in Mosul, Kirkuk, and surrounding areas — has been in centuries-long contact with Kurdish and shows systematic Kurdish influence in phonology, lexis, and morphology that is absent from southern and central Iraqi Arabic. Kurdish loanwords in Mosuli Arabic are more numerous and more deeply integrated than the occasional Kurdish borrowing found in Baghdadi Arabic. Kurdish phonemes — including sounds absent from Arabic — appear in Mosuli Arabic in words of Kurdish origin and in some phonological environments in native Mosuli Arabic vocabulary. Mosuli Arabic also shows Kurdish-influenced morphological patterns not found in other Iraqi sub-dialects.

For DID models, Kurdish-contact features in Mosuli Arabic are a strong positive identifier for northern Iraqi Arabic — but only if the DID training data correctly labels northern Iraqi Arabic as such rather than classifying it as Kurdish-influenced non-Arabic. Standard Arabic DID models, encountering Kurdish-origin vocabulary or phonological patterns in Mosuli Arabic text or speech, have no representation for these features and may classify northern Iraqi Arabic as non-Arabic, as unidentified dialect, or as a lower-confidence Iraqi Arabic determination. Native Mosuli annotators reliably identify Kurdish-contact features as Northern Iraqi Arabic dialect markers; Arabic-only annotators from Baghdad or southern Iraq may not recognise Kurdish-origin vocabulary as part of the Iraqi Arabic lexical system.

4. Urban-rural spectrum within each Iraqi sub-dialect

Each of the three major Iraqi Arabic sub-dialects spans an urban-rural spectrum, with urban varieties (Baghdad city, Basra city, Mosul city) showing different phonological and lexical features from the rural varieties of the same dialect region. Urban Baghdadi Arabic shows more MSA influence and more English loanword integration than rural central Iraqi Arabic. Rural Basrawi Arabic preserves phonological features that have been reduced in urban Basra. The urban-rural axis within each Iraqi sub-dialect is an additional dimension of variation that DID annotation must capture if the resulting model is to distinguish not just the three broad Iraqi sub-dialects but also the within-sub-dialect variation that affects downstream NLP performance.

For production DID applications — routing conversational AI to the appropriate dialect-specialised model, personalising content recommendations by dialect, or segmenting analytics by dialect region — the urban-rural dimension within Iraqi Arabic is often as commercially relevant as the Baghdadi-vs-Basrawi distinction. Native Iraqi annotators who can reliably place text or speech on the urban-rural spectrum within each sub-dialect provide the training data for models that can make this finer-grained classification.

5. Diaspora Iraqi Arabic under contact conditions

Iraqi diaspora communities in Jordan, the UAE, Sweden, Germany, the United Kingdom, and the United States produce Arabic that shows significant accommodation to the surrounding Arabic dialect or European language environment. Diaspora Baghdadi in Jordan shows progressive Levantine phonological and lexical accommodation over residence duration; diaspora Iraqi in the UAE acquires Gulf Arabic loanwords and code-switches into Gulf Arabic for professional and commercial registers. In Western diaspora contexts, Iraqi Arabic shows English code-switching patterns and lexical borrowing that are not found in in-country Iraqi Arabic.

Standard Arabic DID models trained on in-country Iraqi Arabic data encounter diaspora Iraqi Arabic — which retains Mesopotamian core features under contact influence — and either classify it as the contact dialect (Levantine, Gulf) or produce low-confidence Iraqi identification that underestimates the Iraqi component. For platforms with significant Iraqi diaspora user bases — which include many major pan-Arab media, social, and e-commerce platforms operating in Jordan, UAE, and Western markets — accurate DID annotation of diaspora Iraqi Arabic requires explicit labelling protocols for contact-variety identification and native Iraqi annotators who can recognise Mesopotamian features persisting under contact.

The Iraqi DID Training Data Gap

Building an Iraqi Arabic DID system that achieves production-grade sub-dialect accuracy requires a training corpus that explicitly represents all three Iraqi sub-dialects, the urban-rural spectrum within each, and diaspora Iraqi Arabic under various contact conditions. No such corpus exists in publicly available DID training resources. The MADAR corpus Iraqi samples are primarily urban Baghdadi Arabic; Basrawi and Mosuli varieties are marginally represented; diaspora Iraqi Arabic has no systematic coverage in any published Arabic DID corpus.

Our Arabic NLP annotation service provides Iraqi Arabic DID annotation with native annotators across all three major Iraqi sub-dialects and diaspora Iraqi Arabic. Annotation protocols include phonological feature checklists, sub-dialect lexical reference lists, urban-rural spectrum guidance, and diaspora-contact-variety classification procedures.

Need Iraqi Arabic dialect identification annotation?

AI Taggers provides Arabic dialect annotation with native Iraqi annotators. Baghdadi, Basrawi, and Mosuli sub-dialect coverage. Urban-rural spectrum annotation. Diaspora Iraqi Arabic contact-variety labelling.

Get a quote

Case Study: Pan-Arab Streaming Platform — Iraqi Sub-Dialect Accuracy From 54.1% to 86.3%

A pan-Arab streaming and on-demand content platform — serving 18.4 million monthly active users across 22 countries, with Iraqi users comprising 12.3% of total audience — deployed an Arabic dialect identification system to power dialect-aware content recommendations, subtitle dialect matching, and Arabic language analytics. The DID system used a commercial pan-Arabic dialect classifier trained on the MADAR corpus and supplementary Gulf and Egyptian Arabic data.

Before: The DID system achieved 54.1% accuracy at Iraqi sub-dialect level on a held-out evaluation set of 3,800 Iraqi user interaction samples, labelled by native Iraqi annotators. Baghdadi vs Basrawi discrimination accuracy was 61.3% — meaning Basrawi users were classified as Baghdadi in 38.7% of cases and routed to Baghdadi-dialect content rather than Basrawi-relevant recommendations. Mosuli users were classified as Iraqi Arabic in only 48.9% of cases; in 31.2% of cases they were misclassified as Levantine Arabic due to Kurdish phonological features being associated with Levantine Arabic patterns by the DID model. Diaspora Iraqi users in Jordan and UAE were classified as Iraqi in only 39.4% of cases, with the majority classified as Levantine (Jordan diaspora) or Gulf (UAE diaspora).

The practical consequence was that Iraqi sub-dialect content recommendations were delivering the wrong dialect-variant content to 40–61% of Iraqi users, depending on their sub-dialect. Basrawi users received Baghdadi-dialect programming rather than Basrawi-dialect recommendations; Mosuli users received Levantine Arabic content rather than northern Iraqi Arabic-relevant programming; diaspora Iraqi users in the UAE received Gulf Arabic content rather than Iraqi-relevant recommendations. Iraqi user session length was 23.4% shorter than Egyptian and Gulf Arabic user session length on a per-user-per-day basis, a gap the product team attributed in part to dialect-recommendation mismatch.

The DID annotation project delivered 61,000 labelled Iraqi Arabic samples across all three major sub-dialects, the urban-rural spectrum within each, and diaspora Iraqi Arabic in Jordan, UAE, and Western-country contexts. Annotation was conducted by sixteen native Iraqi annotators — seven Baghdadi-native (four urban, three rural-regional), five Basrawi-native, and four Mosuli-native — with supplementary annotation by four Iraqi diaspora-background annotators for the diaspora contact-variety samples. Annotation protocol included phonological feature checklists per sub-dialect, Mosuli Kurdish-contact feature reference, and diaspora accommodation indicators for each contact-variety context. Inter-annotator agreement on the final dataset was 89.2% kappa.

After fine-tuning the DID system on the Iraqi annotation dataset: Overall Iraqi sub-dialect accuracy on the held-out evaluation set rose from 54.1% to 86.3%. Baghdadi vs Basrawi discrimination accuracy improved from 61.3% to 91.7%. Mosuli identification accuracy improved from 48.9% to 83.4%, with Levantine misclassification falling from 31.2% to 6.8%. Diaspora Iraqi identification improved from 39.4% to 74.8% — lower than in-country Iraqi accuracy, reflecting the genuine difficulty of contact-variety classification, but well above the previous threshold at which the DID system was functionally useless for diaspora users.

Iraqi user session length increased 19.7% in the 90 days following the DID system update, narrowing the gap with Egyptian and Gulf Arabic session length from 23.4% to 7.1%. Iraqi user 30-day retention improved 11.3 percentage points. The platform attributed approximately AUD $3.1M in incremental annual subscription revenue to the improved Iraqi user engagement metrics, driven primarily by better dialect-relevant content routing for Basrawi and Mosuli users who had been receiving systematically wrong dialect recommendations under the previous DID system. The annotation project cost AUD $44,000 for annotation, protocol development, inter-annotator agreement measurement, and delivery in the DID fine-tuning format.

What Iraqi Arabic Dialect ID Annotation Requires

Effective Iraqi Arabic DID annotation requires five protocol elements that generic Arabic dialect annotation guidelines do not include.

Phonological feature checklists per sub-dialect. Annotation guidelines must provide annotators with specific phonological features to listen or look for in each Iraqi sub-dialect — Baghdadi gaf, Basrawi emphatic-environment vowel quality, Mosuli Kurdish-contact phonemes — as positive identification criteria. Without explicit feature checklists, native Iraqi annotators apply intuitive holistic judgement that produces adequate IAA at the broad Iraqi vs non-Iraqi level but degrades to 71–78% IAA at the sub-dialect level. Explicit feature checklists bring IAA on sub-dialect classification up to 87–92%.

Sub-dialect lexical reference lists. DID annotation guidelines must include reference lists of lexical items that are sub-dialect-specific — words that appear primarily or exclusively in Baghdadi, Basrawi, or Mosuli Arabic — as positive identification signals for written text DID. These lists cannot be derived from Arabic dialect dictionaries or MSA resources; they must be compiled from native Iraqi speaker knowledge and validated across multiple sub-dialect native speakers before use as annotation reference material.

Urban-rural spectrum classification guidance. Annotation guidelines must specify how annotators should classify text or speech that shows mixed urban and rural features — the case for Iraqi speakers who have moved between rural and urban environments, or for second-generation urban speakers with rural-origin family dialect features. A classification schema that treats the urban-rural spectrum as a continuous dimension rather than a binary produces DID training data that better represents the actual variation in the Iraqi Arabic-speaking population.

Diaspora accommodation indicators. For platforms with diaspora Iraqi user bases, annotation guidelines must specify how to classify Iraqi Arabic that shows partial accommodation to a contact dialect — Gulf, Levantine, or European language — at different stages of accommodation. Classification as "diaspora Baghdadi under Gulf contact" rather than simply "Gulf" or "Iraqi" preserves the sub-dialect information that downstream dialect-aware systems can use for personalisation and analytics.

Multi-annotator adjudication for ambiguous cases. Iraqi Arabic DID at sub-dialect level involves genuine ambiguity for speakers who code-mix across sub-dialects or show contact features from multiple dialect sources. Annotation guidelines must specify adjudication procedures — two-annotator agreement threshold for confidence-weighted labels, three-annotator adjudication for low-agreement cases, quality-assurance sampling rate — that produce DID training data with calibrated confidence scores for ambiguous cases rather than forcing binary classification where the feature evidence is genuinely mixed. Our Arabic dialect annotation service includes all five protocol elements as standard, with delivery validated against 8% blind-sample QA.

Related Reading

Frequently Asked Questions

What is Iraqi Arabic dialect identification annotation?+
Iraqi Arabic DID annotation is the labelling of Mesopotamian Arabic text or speech with its specific sub-dialect origin — Baghdadi, Basrawi, Mosuli, or diaspora Iraqi — by native Iraqi annotators. Standard Arabic DID models achieve only 48–61% accuracy at Iraqi sub-dialect level because the phonological and lexical features distinguishing Iraqi sub-dialects — Baghdadi gaf, Basrawi vowel quality, Kurdish-contact Mosuli features, and diaspora accommodation patterns — are absent from Arabic DID training corpora. Native Iraqi annotators are required to build training data that closes this gap.
Why is Iraqi Arabic the hardest major dialect for Arabic DID models?+
Iraqi Mesopotamian Arabic is structurally the most internally diverse of the major Arabic dialect groups, with three distinct sub-dialects (Baghdadi, Basrawi, Mosuli) that differ from each other in phonology, lexis, and morphology more than internal variation within Gulf, Egyptian, or Levantine Arabic. Additionally, Iraqi Arabic varieties are underrepresented in DID training corpora by a factor of 6–9 compared to Egyptian and Gulf Arabic. Models trained on these corpora identify broad Mesopotamian Arabic but cannot reliably perform sub-dialect classification within the Iraqi group.
Can Arabic DID models distinguish Baghdadi from Basrawi Arabic?+
Standard Arabic DID models achieve Baghdadi vs Basrawi discrimination accuracy of only 52–67%, confusing the two sub-dialects in nearly one-third to half of cases. The Baghdadi qaf-to-gaf phoneme substitution is a strong discriminating feature — categorically present in Baghdadi and absent in Basrawi — but DID models without Iraqi-specific training data have not learned to use this feature consistently. Models without native-annotator Iraqi DID training data default to classifying ambiguous Mesopotamian text as Baghdadi, producing systematic Basrawi misclassification.
Does diaspora Iraqi Arabic affect DID model accuracy?+
Yes. Diaspora Iraqi communities in Jordan, UAE, Sweden, and the UK produce Arabic showing accommodation to the surrounding dialect or language, creating contact varieties that differ systematically from in-country Iraqi Arabic. Standard DID models trained on in-country data misclassify diaspora Iraqi as the contact dialect (Levantine, Gulf) rather than as Iraqi Arabic under contact conditions. Diaspora Baghdadi in Jordan achieves only 39.4% correct Iraqi identification from standard DID systems.
How much does Iraqi Arabic dialect identification annotation cost?+
Native Iraqi-speaker DID annotation for text costs AUD $0.12–$0.28 per item for binary Iraqi vs non-Iraqi classification, and AUD $0.22–$0.48 per item for sub-dialect classification. Speech DID annotation costs AUD $32–$75 per audio hour for sub-dialect classification. Non-native annotation of Iraqi DID produces accuracy of 61–74% on sub-dialect tasks — a gap of 19–29 accuracy points versus native-annotator annotation that is material for production DID systems used in content routing or dialect-aware NLP pipelines.
What annotation protocol is needed for Iraqi Arabic dialect identification?+
Iraqi Arabic DID annotation requires five protocol elements: phonological feature checklists per sub-dialect (Baghdadi gaf, Basrawi vowel quality, Mosuli Kurdish-contact features); sub-dialect lexical reference lists; urban-rural spectrum classification guidance; diaspora accommodation indicators specifying how to label Iraqi Arabic under contact with a second dialect; and multi-annotator adjudication procedures for ambiguous cases. Inter-annotator agreement on Iraqi Arabic sub-dialect DID using these protocol elements reaches 87–92% kappa, compared to 63–71% kappa using generic Arabic dialect guidelines applied by non-Iraqi annotators.
Free Sample · 24-48 hours

Get a Quote for Iraqi Arabic Dialect Identification Annotation

Native Iraqi annotators. Baghdadi, Basrawi, and Mosuli sub-dialect coverage. Urban-rural spectrum annotation. Diaspora Iraqi Arabic contact-variety labelling for Jordan, UAE, and Western contexts.

No commitment. NDA available on request. We respond within 24 hours, often the same day for Gulf-region inquiries.

Neel Bennett

AI Annotation Specialist at AI Taggers

Neel has over 8 years of experience in AI training data and machine learning operations. He specializes in helping enterprises build high-quality datasets for computer vision and NLP applications across healthcare, automotive, and retail industries.

Connect on LinkedIn