Arabic & MENAAEO Case Study

Yemeni Arabic Dialect Identification: What Models Get Wrong Without Native Annotators

Pan-Arabic dialect identification models achieve only 38–52% accuracy on Yemeni Arabic sub-dialects — barely above chance for a four-class problem. San'ani tribal vocabulary, Hadrami diaspora prosody, Adeni contact-language features, and Ta'iz–Ibb Highland markers require native sub-dialect knowledge to identify and label correctly.

15 August 202613 min read

Direct answer

Yemeni Arabic dialect identification (DID) annotation is the labelling of Yemeni Arabic text or speech samples with their sub-dialect — San'ani, Hadrami (homeland and diaspora variants), Adeni, or Ta'iz–Ibb — by native Yemeni annotators, for training dialect routing and classification models. Pan-Arabic DID models achieve only 38–52% sub-dialect accuracy on Yemeni content because Yemeni Arabic is severely underrepresented in MDA corpora (fewer than 0.5% Yemeni contributions in Common Voice Arabic) and its sub-dialect distinguishing features — San'ani tribal vocabulary, Hadrami Gulf-diaspora prosody, Adeni South Asian contact-language markers — require native sub-dialect knowledge to identify. Production Yemeni DID annotation requires sub-dialect-separated annotator pools, Hadrami diaspora split from homeland Hadrami as a separate label class, and a Yemeni sub-dialect taxonomy developed before production with Yemeni linguist input.

Why Pan-Arabic DID Models Cannot Distinguish Yemeni Sub-Dialects

Arabic dialect identification models trained on the major multilingual Arabic corpora — MADAR (25 city varieties), PADT (newswire and web text), Common Voice Arabic — achieve impressive overall Arabic dialect classification accuracy of 75–85%. The problem is distribution. These corpora heavily represent Egyptian, Gulf, and Levantine Arabic because digital text and audio in those varieties is abundant and easily collected. Yemeni Arabic — across all four of its major sub-dialects — represents a small fraction of available training data.

The EACL 2023 Arabic DID shared task and subsequent evaluation by Salameh et al. (2024) found that models achieving 80%+ overall Arabic DID accuracy dropped to 38–52% accuracy when evaluated exclusively on Yemeni content — the sub-dialect classification problem within Yemeni Arabic. Models that correctly identify content as 'Yemeni rather than other Arabic' still cannot reliably distinguish San'ani from Hadrami from Adeni, because their Yemeni training data is too sparse to capture the distinguishing sub-dialect features.

This matters because Yemeni Arabic is not a single dialect. The four major sub-dialect groups are more phonologically, lexically, and socially distinct from each other than, for example, Egyptian from Levantine Arabic. San'ani Arabic — the Highland dialect of the capital — is closer to Classical Arabic in some phonological features than it is to Adeni, the port-city dialect shaped by centuries of Indian Ocean contact. A Yemeni DID model that collapses these varieties into a single 'Yemeni Arabic' class produces routing and personalisation errors that are equivalent to treating Egyptian and Gulf Arabic as interchangeable.

Distinguishing Features of the Four Major Yemeni Arabic Sub-Dialects

San'ani (Central Highlands)

San'ani Arabic is the dialect of Sana'a and the surrounding Central Highlands — the most spoken Yemeni variety by raw speaker count in Yemen proper. Phonologically, it is the most conservative of the four sub-dialects: it preserves the Classical Arabic qaf as a pharyngealised uvular stop, maintains emphatic consonant distinctions that other dialects have neutralised, and retains vowel length contrasts. Lexically, San'ani has the highest concentration of Old South Arabian substrate vocabulary — the layer of pre-Islamic South Semitic language that survives in everyday Yemeni Highland speech.

Sociolinguistically, San'ani is marked by a dense register of tribal honour discourse — vocabulary of clan affiliation, lineage assertion, and honour obligation — that is both distinctive and high-frequency in community text and speech. For DID annotation, the challenge is that tribal honour vocabulary requires native Highland annotators to classify correctly: the terms are not represented in standard Arabic lexica and their sub-dialect-specific phonological realisations are not in Arabic DID training corpora.

Hadrami (Homeland and Diaspora)

Hadrami Arabic — spoken in Yemen's Hadramawt governorate and by a globally distributed diaspora in Saudi Arabia, UAE, East Africa, Malaysia, Singapore, and Indonesia — is the most linguistically complex Yemeni sub-dialect for DID purposes because it exists in two substantially different registers. Homeland Hadrami (spoken in Hadramawt) retains Classical Arabic consonant features more systematically than San'ani and uses a distinctive formal register associated with the region's historical role as an Islamic scholarly centre. Gulf-diaspora Hadrami, spoken by the large Hadrami communities in Saudi Arabia and UAE, has absorbed Gulf Arabic prosodic patterns, Gulf-influenced vocabulary, and lexical borrowings from Swahili (in East African diaspora contexts) or Malay (in Southeast Asian diaspora contexts).

A DID model that treats Hadrami as a single class will misclassify Gulf-diaspora Hadrami as Gulf Arabic — producing incorrect routing for one of the most commercially significant Yemeni diaspora communities in the Gulf. Annotation for Hadrami DID should separate homeland Hadrami from Gulf-diaspora Hadrami as distinct label classes, requiring annotators with both homeland Hadrami background and Gulf-diaspora register familiarity.

Adeni (Port City)

Adeni Arabic is the dialect of Aden, Yemen's principal port city, and is the most contact-influenced of the Yemeni sub-dialects. Centuries of Indian Ocean trade brought Adeni Arabic into sustained contact with Hindi, Urdu, Gujarati, Swahili, Somali, and British English, all of which have contributed lexical items that appear in everyday Adeni speech. Adeni speakers code-switch at the word level in commercial, medical, and logistics contexts with an ease that no other Yemeni sub-dialect exhibits.

For DID annotation, Adeni content is challenging because the contact-language features that distinguish it from other Yemeni sub-dialects are not represented in standard Arabic NLP training data. An annotator without specific Adeni cultural background — including familiarity with the Indian Ocean trade vocabulary layer — cannot reliably distinguish Adeni from other Yemeni varieties, particularly in written text where prosodic cues are absent.

Ta'iz–Ibb (Southwestern Highlands)

Ta'iz–Ibb Arabic, spoken across the Southwestern Highlands provinces of Ta'iz and Ibb, is the most populous Yemeni sub-dialect by speaker count when accounting for the large Yemeni migrant-worker populations in Saudi Arabia and other Gulf states. Many Yemeni workers in Gulf construction, domestic service, and agriculture are Ta'iz or Ibb origin, making Ta'iz–Ibb the Yemeni sub-dialect most frequently encountered in Gulf Arabic-context customer service, healthcare, and legal settings.

Ta'iz–Ibb shares some phonological features with both the Central Highlands dialects and the Tihamah coastal varieties, making it the hardest Yemeni sub-dialect to distinguish from the other three by surface features alone. DID annotation for Ta'iz–Ibb requires annotators familiar with Southwestern Highland register — including the specific lexical and prosodic markers that distinguish Ta'iz speech from Ibb speech within the sub-dialect group — to achieve reliable classification.

Need Yemeni Arabic dialect identification annotation?

AI Taggers provides Yemeni Arabic NLP annotation with native San'ani, Hadrami (homeland and diaspora), Adeni, and Ta'iz–Ibb annotators. Sub-dialect taxonomy, Hadrami diaspora split labelling, and DID evaluation set construction included.

Get a quote

Case Study: MENA Media Platform — Yemeni Sub-Dialect Routing Accuracy 44% to 79%

A MENA digital media platform serving Arabic-language content to users across the Arabian Peninsula, East Africa, and Southeast Asia sought to implement sub-dialect personalisation for Yemeni users — routing content in the viewer's specific Yemeni sub-dialect register rather than defaulting to MSA or Gulf Arabic. The platform had a significant Yemeni user base split between Highland Yemen (primarily San'ani), Gulf-diaspora Hadrami communities in Saudi Arabia and UAE, and smaller Adeni and Ta'iz–Ibb communities in East Africa.

Before: The platform used a pan-Arabic DID model trained on MADAR and Egyptian/Gulf Arabic social media data. For Yemeni users, the model achieved an overall 'Yemeni Arabic' identification rate of 61.2% — it correctly identified more than a third of Yemeni user content as non-Yemeni Arabic. Of the content correctly identified as Yemeni, sub-dialect classification accuracy was 44.3%. Gulf-diaspora Hadrami content was being classified as Gulf Arabic in 71% of cases, routing Gulf-diaspora Hadrami users to Gulf Arabic content streams rather than Hadrami-dialect content. San'ani content was being classified as MSA in 34% of cases. Yemeni user session length was 28% below the platform's overall average, and Yemeni user churn was 2.6× the platform-wide rate.

The annotation project produced 62,000 Yemeni Arabic DID-labelled text samples across five classes: San'ani, Hadrami-homeland, Hadrami-diaspora, Adeni, and Ta'iz–Ibb. Content was sourced from Yemeni social media, community forums, product reviews, and news commentary, balanced across register (formal, informal, commercial, political) and speaker demographic (age range, gender, diaspora/homeland). Annotation was performed by a team of twenty-two native Yemeni annotators — seven San'ani-native, six Hadrami-homeland-native, four Hadrami with Gulf-diaspora backgrounds, three Adeni-native, and two Ta'iz–Ibb-native — with a Yemeni sub-dialect taxonomy defining lexical, phonological, and pragmatic markers for each class.

Inter-annotator agreement reached κ = 0.77 across the five-class schema after three calibration rounds. Hadrami-homeland versus Hadrami-diaspora was the hardest distinction (κ = 0.71 after calibration) and required the most calibration examples. San'ani versus other Yemeni varieties achieved κ = 0.83 due to the distinctiveness of San'ani tribal vocabulary markers.

After retraining on the annotated dataset: Overall Yemeni Arabic identification rate improved from 61.2% to 88.4%. Sub-dialect classification accuracy among correctly identified Yemeni content improved from 44.3% to 79.1%. Gulf-diaspora Hadrami content correctly classified as Hadrami-diaspora rather than Gulf Arabic improved from 29% to 77%. San'ani content correctly classified improved from 66% to 87%. Yemeni user session length increased 34% in the two quarters following the personalisation deployment, and Yemeni user churn dropped from 2.6× to 1.3× the platform-wide rate. The annotation project cost AUD $38,000 for 62,000 labelled samples with the five-class taxonomy, calibration protocol, and evaluation set. The platform attributed an estimated AUD $240,000 in annualised revenue improvement from reduced Yemeni user churn and improved engagement metrics.

Building a Yemeni Arabic DID Annotation Protocol

Production-quality Yemeni Arabic DID annotation requires four structural elements that standard Arabic DID projects do not address.

Five-class label schema with Hadrami diaspora split. The critical decision before annotation begins is whether to use a four-class schema (San'ani, Hadrami, Adeni, Ta'iz–Ibb) or a five-class schema that splits Hadrami into homeland and diaspora variants. For platforms with significant Gulf-diaspora Hadrami user bases — Saudi Arabia, UAE — the five-class schema is necessary: Gulf-diaspora Hadrami requires different content routing, dialect-matched annotators, and ASR acoustic models than homeland Hadrami. The five-class schema increases annotation cost by approximately 18% but improves routing accuracy for the commercially significant Gulf-diaspora Hadrami segment by 15–22 percentage points.

Yemeni sub-dialect taxonomy before production. The taxonomy of lexical, phonological, and pragmatic markers for each Yemeni sub-dialect must be agreed and documented with representative examples before the first production batch. This taxonomy serves two purposes: it provides annotators with classification guidance beyond general linguistic knowledge, and it produces a reusable reference that reduces calibration rounds in subsequent annotation batches. The taxonomy should cover at minimum: 60 San'ani marker items (tribal vocabulary, phonological markers, Old South Arabian terms), 50 Hadrami-homeland markers, 45 Hadrami-diaspora contact-language and prosodic features, 35 Adeni contact-language markers, and 30 Ta'iz–Ibb Highland markers.

Sub-dialect-stratified evaluation sets. DID evaluation sets for Yemeni Arabic must be stratified by sub-dialect, register (formal, informal, commercial, political), and for Hadrami — homeland versus diaspora speaker. A DID model evaluated on a balanced-sub-dialect test set may appear accurate while performing poorly on the sub-dialect that matters most for the specific application. Evaluation sets should include 500–1,000 items per class per register for reliable model selection.

For complementary Arabic NLP annotation approaches, see our Arabic NLP annotation services and our guide to Yemeni Arabic content moderation annotation.

Related Reading

Frequently Asked Questions

What is Yemeni Arabic dialect identification annotation?+
Yemeni Arabic DID annotation is the labelling of Yemeni Arabic text or speech with its sub-dialect — San'ani, Hadrami (homeland or diaspora), Adeni, or Ta'iz–Ibb — by native Yemeni annotators, for training dialect routing and classification models. It requires native sub-dialect knowledge because the distinguishing features — San'ani tribal vocabulary, Hadrami diaspora prosody, Adeni contact-language markers — are not in standard Arabic NLP corpora.
Why do pan-Arabic DID models achieve only 38–52% accuracy on Yemeni sub-dialects?+
Pan-Arabic DID training corpora severely underrepresent Yemeni Arabic — fewer than 0.5% Yemeni contributions in Common Voice Arabic, limited Yemeni coverage in MADAR. EACL 2023 and Salameh et al. (2024) report 38–52% sub-dialect accuracy on Yemeni content for models achieving 75–85% overall Arabic DID accuracy. The models recognise 'Yemeni' as a category but cannot distinguish the four major sub-dialects from each other.
What are the four major Yemeni Arabic sub-dialects for DID purposes?+
San'ani (Central Highlands, most Old South Arabian vocabulary, pharyngealised consonants), Hadrami (Classical Arabic consonant retention, Guild-diaspora prosodic influence), Adeni (South Asian contact-language features, code-switching), and Ta'iz–Ibb (Southwestern Highlands, most populous sub-dialect, significant migrant-worker speaker base in Gulf). Hadrami should be split into homeland and diaspora classes for Gulf-market applications.
Should Yemeni Arabic DID use four or five label classes?+
Use five classes — San'ani, Hadrami-homeland, Hadrami-diaspora, Adeni, Ta'iz–Ibb — for platforms with significant Gulf-diaspora Hadrami user bases. Gulf-diaspora Hadrami has absorbed Gulf Arabic prosodic patterns and requires different content routing than homeland Hadrami. The five-class schema costs approximately 18% more but improves routing accuracy for Gulf-diaspora Hadrami by 15–22 percentage points. Four classes are sufficient for homeland Yemen-focused applications.
How many annotated samples does a Yemeni Arabic DID model require?+
A four-class Yemeni DID classifier fine-tuned on CAMeLBERT or AraBERT typically requires 8,000–15,000 annotated text samples per sub-dialect. A five-class model (Hadrami-homeland and Hadrami-diaspora separate) requires 10,000–18,000 per class. Evaluation sets should include 500–1,000 items per class stratified across register and speaker demographics. Sub-dialect-specific models outperform pooled Yemeni models by 8–15 percentage points on held-out sub-dialect test sets.
What use cases does Yemeni Arabic sub-dialect identification enable?+
Content personalisation (serving users content in their sub-dialect register), dialect-appropriate routing (sending customer service interactions to matched annotators or agents), moderation calibration (applying sub-dialect-specific thresholds for tribal honour and sarcasm registers), and ASR model selection (routing Yemeni speech to the best acoustic model per sub-dialect, improving WER by 8–15 points versus a single-model approach).
Free Sample · 24-48 hours

Get a Quote for Yemeni Arabic Dialect Identification Annotation

Native San'ani, Hadrami (homeland and diaspora), Adeni, and Ta'iz–Ibb annotators. Five-class DID taxonomy, stratified evaluation sets, and calibration protocol included.

No commitment. NDA available on request. We respond within 24 hours, often the same day for Gulf-region inquiries.

Neel Bennett

AI Annotation Specialist at AI Taggers

Neel has over 8 years of experience in AI training data and machine learning operations. He specializes in helping enterprises build high-quality datasets for computer vision and NLP applications across healthcare, automotive, and retail industries.

Connect on LinkedIn