Arabic & MENAAEO Case Study

Levantine Arabic Named Entity Recognition: What Models Get Wrong Without Native Annotators

Standard Arabic NER models lose 24–37% F1 on Levantine entities. Shami person names with Aramaic roots, Beirut neighbourhood references, Lebanese French-influenced organisation names, and code-switched entity spans break every standard Arabic NER pipeline. Here is the native-speaker annotation approach that fixes it.

1 August 202613 min read

Direct answer

Levantine Arabic NER annotation is the identification and classification of named entities — persons, organisations, locations, dates, and domain-specific items — in Shami-dialect text by native-speaker annotators. Standard Arabic NER models lose 24–37% F1 on Levantine test sets because Shami Arabic has Syrian and Lebanese person name morphology absent from standard gazetteers, Beirut and Damascus neighbourhood references not in MSA location corpora, French-influenced Lebanese organisation names, and code-switched entity spans where entity names appear in French or English within Arabic text. Effective Levantine NER annotation requires dialect-native annotators from Syria and Lebanon, bilingual handling for French-Arabic code-switched text, domain-entity supplements for financial and government applications, and appropriate compliance treatment for personal name data under Lebanese and Jordanian privacy law.

Why Standard Arabic NER Breaks on Levantine Text

Named entity recognition for Arabic is a solved problem — if your text is MSA newswire from Gulf or Egyptian sources. The moment your production corpus shifts to Levantine dialect, the problem becomes substantially harder. Syrian social media, Lebanese business communications, Palestinian NGO documents, and Jordanian government services text all contain named entity patterns that are systematically absent from Arabic NER training data, causing the entity-level precision and recall failures that cascade into document processing and information extraction failures downstream.

The Arabic NER corpora that underpin most production models — ANERCorp (500 MSA news articles), AQMAR (Wikipedia-derived), and the CoNLL-2012 Arabic OntoNotes subset — were built from formal MSA sources. Person name gazetteers in these corpora reflect Egyptian, Gulf, and Levantine names proportionally to their representation in pan-Arab media, which substantially under-represents distinctly Levantine name morphology. Location gazetteers cover Cairo, Riyadh, Dubai, and pan-Arab capitals but have minimal coverage of Beirut districts, Damascus neighbourhoods, or Syrian city administrative subdivisions.

Research quantifying this gap (Shahrour et al., Arabic NLP Workshop 2021; Jarrar et al., 2022) shows that Arabic NER models trained on standard corpora achieve 24–37% lower F1 scores on Levantine-dialect test sets compared to their performance on MSA test sets. The F1 degradation is not uniform: person entity recall drops most sharply (often 30–45% degradation), followed by organisation recall (20–35%), with location entity performance showing the widest variance depending on whether the location is a major city or a sub-city district.

Five Levantine Entity Categories That Break Standard NER

1. Aramaic-origin and Christian Lebanese given names

Lebanese person names include a substantial proportion of names with Aramaic, Syriac, or French roots that are absent from MSA and Egyptian NER gazetteers. Names like Jad, Tala, Celine, Charbel, Rima, Mirna, and Nour are common in Lebanese and Syrian Christian communities but appear rarely in Egyptian and Gulf Arabic NER training data. A standard Arabic NER model encountering ‘Jad Khalil’ in a Lebanese support ticket will frequently fail to tag it as a PERSON entity because neither name matches the model's Arabic name patterns.

The scale of this gap in Lebanese commercial text is substantial: a 2023 audit of Lebanese fintech customer data found that 38% of person entity spans in support tickets contained at least one name component from the under-represented category — Aramaic-origin given name, French family name, or French-Arabic hyphenated name. Standard NER models correctly tagged only 54% of these spans, compared to 91% tagging accuracy for Egyptian-pattern Arabic names in the same corpus.

2. Beirut district and Damascus neighbourhood names

Levantine location references in commercial and social text frequently use sub-city district names and colloquial neighbourhood references that are absent from standard Arabic NER location gazetteers. Beirut districts — Mar Mikhael, Gemmayzeh, Hamra, Achrafieh, Badaro, Verdun, Jdeideh — are well-known to Lebanese users but absent from MSA location corpora. Syrian neighbourhood names within Damascus (Mezzeh, Kafr Sousa, Malki, Abu Rummaneh) and Aleppo (Hamdaniyeh, Aziziyeh, Sha'ar) present the same coverage gap.

For location-dependent applications — logistics, ride-sharing, food delivery, real estate — missing these district names causes entity misclassification that breaks the entire downstream geocoding and routing pipeline. A delivery address that references ‘Mar Mikhael, Beirut’ is useless if the NER model tags ‘Mar’ as a misrecognised preposition and leaves ‘Mikhael’ untagged.

3. French-influenced Lebanese organisation names

Lebanese organisations — banks, law firms, medical clinics, media companies, NGOs — frequently have French names or French-Arabic hybrid names. ‘Banque du Liban’ (central bank), ‘Byblos Bank’, ‘Clinique du Levant’, and the many French-named civil society organisations appear in Lebanese Arabic text as both full French names and Arabic transliterations. Standard Arabic NER models that lack French entity coverage mis-tag French organisation names as generic nouns or fail to recognise them as entities entirely.

This is particularly problematic for Lebanese financial services NLP — loan processing, KYC document analysis, and contract extraction — where organisation name accuracy is a regulatory compliance requirement, not just a quality metric.

4. Code-switched entity spans

Lebanese Arabic text regularly embeds entity names in French or English within an Arabic sentence structure. A customer support message might read: “شحنت من Aramex بس ما وصلت للـ Mar Mikhael branch” (“I shipped through Aramex but it didn't arrive at the Mar Mikhael branch”). The organisation entity ‘Aramex’ and the location entity ‘Mar Mikhael branch’ both appear in Latin script within an Arabic sentence. Standard Arabic NER models — which operate on Arabic script tokens — either skip these spans entirely or produce fragmented entity boundaries around the code-switch.

5. Palestinian and Syrian conflict-related entity patterns

Palestinian and Syrian text in political, humanitarian, and civic contexts contains organisation names, location references, and event entities related to conflict, displacement, and reconstruction that are absent from standard Arabic NER training data. NGO names, camp names, checkpoint names, and organisation abbreviations from the Syrian conflict and Palestinian political landscape form a distinct entity vocabulary that requires native-speaker annotation with contextual knowledge of the relevant conflict and governance structures.

Need Levantine Arabic NER annotation?

AI Taggers provides Levantine Arabic NLP annotation with native Damascene, Beiruti, Palestinian, and Jordanian annotators, bilingual French-Arabic QA, and domain-entity gazetteer supplements.

Get a quote

Case Study: Lebanese Fintech KYC Extraction — 52% to 86% Entity F1

A Lebanese digital bank processing KYC documents and customer correspondence needed to automate entity extraction from Arabic and mixed-language customer text — support messages, identity document descriptions, and address references. Their existing pipeline used an AraBERT-based NER model trained on standard MSA corpora.

Before: The MSA-trained NER model achieved 52.3% entity F1 on held-out Lebanese customer text. Person entity recall stood at 48.7% — with specific failure on Aramaic-origin names (Jad, Charbel, Mirna, Tala) where recall dropped to 31.2%. Location entity F1 was 44.8%, with Beirut district names — Hamra, Mar Mikhael, Ashrafieh, Badaro — correctly tagged only 28.4% of the time. Organisation entity recall was 61.2%, but fell to 38.7% for French-named or French-transliterated Lebanese organisations. Straight-through KYC processing rate was 19% — meaning 81% of documents required manual review, creating a significant operational bottleneck.

The annotation project delivered 31,000 annotated sentences covering person, organisation, location, date, and financial entity types across Lebanese Arabic, Lebanese French-Arabic mixed text, and formal MSA document text. Seven native annotators — three Beiruti bilingual (Arabic-French), two Damascene-native, two Jordanian-native — conducted annotation with domain-entity guidelines covering Lebanese banking organisation names, Beirut district gazetteers, and Aramaic-origin name lists. French-Arabic code-switched entity spans received bilingual adjudication. Final IAA F1 across entity categories was 0.87.

After fine-tuning on the annotated dataset: Overall entity F1 improved from 52.3% to 86.1%. Person entity recall improved from 48.7% to 83.9%, with Aramaic-origin name recall improving from 31.2% to 79.4%. Location entity F1 improved from 44.8% to 84.2%, with Beirut district name recall improving from 28.4% to 81.6%. Organisation entity recall improved from 61.2% to 87.3% overall, and French-named organisation recall improved from 38.7% to 82.1%. Straight-through KYC processing rate increased from 19% to 64%, reducing the manual review backlog by 73%.

The annotation project cost AUD $48,600 for annotation, bilingual QA, gazetteer supplements, and delivery. The bank estimated AUD $1.8M in annual operational cost savings from the 73% reduction in manual KYC review volume, plus improved compliance velocity as straight-through processing reduced average KYC completion time from 4.2 days to 1.1 days.

The Annotation Protocol for Levantine NER Projects

Effective Levantine NER annotation requires a protocol that generic Arabic NLP pipelines cannot deliver. The essential elements are:

Domain gazetteer supplements before annotation begins. Standard Arabic NER gazetteers must be supplemented with Levantine-specific entity lists: Beirut district names, Damascus and Aleppo neighbourhood names, Lebanese and Syrian person name morphology patterns (including Aramaic-origin names), Lebanese organisation names (with French variants), and Palestinian location and organisation references. Building these gazetteers requires native-speaker input before annotation starts — they cannot be derived from existing Arabic NLP resources.

Bilingual annotation for French-Arabic code-switched text. Lebanese text with French entity spans requires annotators fluent in both languages who can identify and correctly span French-language entities within an Arabic sentence. A monolingual Arabic annotator will mis-span or skip these entities. For Lebanese commercial applications — banking, logistics, e-commerce — this bilingual capability is not optional.

Sub-entity type coverage for domain applications. For financial services NER, the standard PERSON/ORG/LOC/DATE schema is insufficient. Financial entity types — account identifiers, financial instrument names, Lebanese banking regulation references — require domain-specific entity classes that must be defined before annotation begins, with example sentences showing annotation of each sub-type.

IAA tracking by entity category, not just overall. Overall IAA kappa masks category-level annotation failures. For Levantine NER, IAA should be tracked separately for person names (where Aramaic-origin names are the hardest category), location names (where district-level names diverge from annotator familiarity), and organisation names (where French-Arabic hybrids produce the most disagreement). Category-level IAA below 0.75 is a signal to revise guidelines or add example sentences before continuing.

AI Taggers’ Levantine Arabic NLP service covers the full Shami dialect cluster with domain gazetteer supplements, bilingual French-Arabic QA, and the sub-dialect annotation routing that production Levantine NER projects require.

Compliance for Levantine NER Projects

NER annotation on Levantine text has an inherent compliance tension: the task requires labelling person names and location references, which are themselves personal data under Lebanese Law No. 81 of 2018 and Jordan's PDPL of 2023. Both laws define personal data to include names, locations, and any information that could identify an individual — meaning the annotated output of a Levantine NER project is, by design, a labelled personal data corpus.

The practical compliance approach for Levantine NER annotation is to work from source documents that have already had direct contact details removed — phone numbers, email addresses, account numbers, and any identifier beyond name and location — before annotation begins. The named entity spans themselves (names, locations, organisations) are necessary for the annotation task and do not need to be removed. The access log, annotator non-disclosure, and data deletion protocol at project close are the primary compliance controls.

For pan-Levantine projects that also involve Gulf Arabic data, see our post on PDPL vs GDPR for annotation vendors and our end-to-end Arabic data labelling pipeline case study for implementation detail on multi-jurisdiction data handling.

Related Reading

Frequently Asked Questions

What is Levantine Arabic named entity recognition?+
Levantine Arabic NER is the identification and classification of named entities — persons, organisations, locations, dates — in Shami-dialect text from Syria, Lebanon, Palestine, and Jordan. It requires native Levantine annotators because Shami Arabic has person name morphology (Aramaic-origin names, French-influenced Lebanese names), location naming conventions (Beirut districts, Damascus neighbourhoods), and French code-switched entity spans absent from standard Arabic NER gazetteers.
Why do standard Arabic NER models fail on Levantine text?+
Standard Arabic NER is trained on MSA newswire with Egyptian and Gulf name patterns. Levantine introduces Aramaic-origin given names (Jad, Charbel, Mirna), Beirut district names (Mar Mikhael, Hamra, Achrafieh), French-influenced Lebanese organisation names, and French-Arabic code-switched entity spans. Research shows 24–37% F1 degradation on Levantine test sets vs MSA benchmarks, with person recall dropping 30–45%.
What entity types are hardest for standard NER on Levantine text?+
Person names — particularly Aramaic-origin Lebanese and Syrian given names — show the sharpest recall drop (30–45% degradation). Beirut district and Damascus neighbourhood location names show the widest variance. French-influenced Lebanese organisation names show 20–35% recall degradation. French-Arabic code-switched entity spans are the most difficult category overall, as they require bilingual annotators to correctly identify and span.
How much Levantine NER training data is needed?+
Fine-tuning on Levantine NER typically requires 8,000–20,000 annotated sentences for general entity types. Domain-specific NER (financial, healthcare, government) adds 4,000–8,000 domain-entity sentences. A 3,000-sentence IAA-calibrated pilot with Damascene and Beiruti annotators is recommended to identify gazetteer gaps before full-scale production.
What compliance rules apply to Levantine NER annotation?+
NER annotation on Levantine text labels person names and locations — both personal data under Lebanese Law No. 81 (2018) and Jordan's PDPL (2023). The practical approach: remove direct contact details (phone, email, account numbers) before annotation; retain named entity spans as required for the task; implement annotator NDA and data deletion protocol at project close.
What does Levantine Arabic NER annotation cost per sentence?+
Native-speaker Levantine NER costs AUD $0.18–$0.45 per sentence for standard four-class NER. French-Arabic code-switched Lebanese text runs AUD $0.30–$0.65 per sentence due to bilingual annotator requirements. Domain-specific NER (financial, healthcare, legal entities) adds 20–35% to standard pricing due to the domain knowledge and gazetteer supplement requirements.
Free Sample · 24-48 hours

Get a Quote for Levantine Arabic NER Annotation

Native Damascene, Beiruti, Palestinian, and Jordanian annotators. Bilingual French-Arabic QA. Domain gazetteer supplements included.

No commitment. NDA available on request. We respond within 24 hours, often the same day for Gulf-region inquiries.

Neel Bennett

AI Annotation Specialist at AI Taggers

Neel has over 8 years of experience in AI training data and machine learning operations. He specializes in helping enterprises build high-quality datasets for computer vision and NLP applications across healthcare, automotive, and retail industries.

Connect on LinkedIn