Arabic & MENAHeritage AI

Classical and Quranic Arabic NLP: Annotation for Religious and Heritage AI

Modern Arabic NLP annotation methods fail on Classical and Quranic Arabic because the language, morphology, and required annotator qualifications are fundamentally different. Here is the specialist annotation methodology for heritage and religious AI — and a case study from a GCC Islamic knowledge platform that quantifies the impact.

19 August 202614 min read

Direct answer

Classical Arabic NLP annotation is the process of labelling pre-modern Arabic texts — including Quranic text, Hadith literature, classical jurisprudence (fiqh), and heritage manuscripts — for morphological analysis, NER, topic classification, and other NLP tasks. It requires annotators with classical Arabic linguistic training and, for theologically significant tasks, Islamic science expertise. Modern Standard Arabic NLP models perform at 45–60% accuracy on Classical Arabic tasks without fine-tuning; qualified-annotator fine-tuning datasets of 5,000–50,000 items raise this to 80–92%, depending on task complexity.

Why Classical Arabic Is a Different NLP Problem From Modern Arabic

Modern Standard Arabic (MSA) — the formal written Arabic used in news, government, and education — descends from Classical Arabic but differs from it substantially. Classical Arabic is the language of the Quran, of pre-Islamic poetry, of the great classical Islamic scholars (Ibn Sina, Al-Ghazali, Ibn Khaldun), of medieval Islamic jurisprudence, and of a rich tradition of history, philosophy, and literature produced over fourteen centuries.

The morphological gap is significant. Classical Arabic retains the full case-marking system (i'rab) that MSA uses in formal contexts but drops informally. Classical verb forms include patterns rare in MSA. The jussive and energetic moods, common in Quranic verse and classical prose, appear infrequently in modern corpora. An NLP model trained on contemporary news Arabic lacks the morphological coverage to parse classical texts accurately.

The vocabulary gap is equally substantial. Classical texts use words and semantic senses that do not appear in modern Arabic — or that appear in modern Arabic with shifted meanings. A classical term for “scholar” may differ from the contemporary term. Legal vocabulary in classical fiqh uses technical definitions that were codified over centuries of jurisprudential debate. An NLP model that has not been trained on classical vocabulary cannot perform NER or topic classification on these texts reliably.

According to research from the Arabic NLP lab at NYU Abu Dhabi (CAMeL Lab), leading Arabic NLP models — including AraBERT and CAMeLBERT — achieve only 45–62% F1 on Classical Arabic morphological tagging tasks in zero-shot conditions, compared to 85–92% F1 on equivalent MSA tasks. The gap is not architectural — it is a training data gap.

The Islamic AI market amplifies this challenge. Across Saudi Arabia, UAE, Qatar, and Malaysia, a wave of Islamic fintech, Islamic knowledge platforms, Quran recitation apps, and Islamic legal AI systems is creating demand for NLP on Classical and Quranic Arabic text at scale. Saudi Arabia's Vision 2030 explicitly includes Islamic heritage digitisation as a cultural objective. The King Abdulaziz Public Library has digitised over 4 million manuscripts. These digitised collections become AI training opportunities only when correctly annotated.

What Quranic Arabic Annotation Involves

The Quran is the most-studied text in Arabic NLP, and the most consequential to annotate correctly. Several open annotation datasets exist, most notably the Quranic Arabic Corpus from the University of Leeds, which provides morphological annotations (part-of-speech, root, stem, features) for every word in the Quran. This corpus is a foundational resource for Quranic NLP research.

However, commercial Quranic AI applications require annotation beyond what open datasets provide:

A critical principle for all Quranic annotation: the annotation must be reviewed by scholars with appropriate religious credentials for any task that involves interpretation of meaning. This is not just a quality requirement — it is a cultural requirement that affects whether the resulting product will be accepted by its intended audience.

Hadith and Classical Islamic Text Annotation

Hadith literature — the collected sayings and actions of the Prophet Muhammad (peace be upon him) as recorded by his companions — is the second major corpus for Islamic AI annotation. The six canonical Hadith collections (Sahih Al-Bukhari, Sahih Muslim, Sunan Abu Dawud, Jami' At-Tirmidhi, Sunan An-Nasa'i, Sunan Ibn Majah) contain over 30,000 individual narrations, each with a chain of transmission (isnad) and a body text (matn).

NLP annotation tasks on Hadith literature include:

Classical Islamic jurisprudence texts (fiqh) — the Hanafi, Maliki, Shafi'i, and Hanbali madhab texts — add another annotation domain with a distinct vocabulary of legal concepts, case reasoning structures, and scholarly commentary. Fiqh NLP is nascent but growing rapidly in the GCC as Islamic fintech requires AI systems that can analyse contracts for Sharia compliance.

Building a Classical or Quranic Arabic AI Dataset?

AI Taggers provides Arabic NLP annotation with annotators qualified in classical Arabic linguistics and Islamic sciences. We provide morphological tagging, Hadith NER, Quranic thematic classification, and tafsir alignment — with scholar review for theologically sensitive tasks.

Discuss your project

Case Study: GCC Islamic Knowledge Platform — Quranic Search Accuracy From 51% to 83%

A GCC-based Islamic knowledge platform serving users across Saudi Arabia, UAE, Kuwait, and Malaysia engaged AI Taggers in early 2026 to improve the quality of their Quran and Hadith search system. The platform had deployed a fine-tuned AraBERT model for semantic search but user satisfaction surveys showed significant frustration with search results — users felt the system was returning passages that were technically relevant by keyword overlap but wrong by Islamic scholarly context.

An audit of 1,000 search query-result pairs found the core issue: the existing training data had been annotated by native Arabic speakers without classical Arabic or Islamic science training. Annotators had correctly identified surface-level semantic relevance but had missed the distinction between passages that address the same topic and passages that address the same legal question — a distinction that matters enormously to users of an Islamic legal research tool.

The annotation project involved two phases:

  1. Quranic thematic annotation: 6,240 verse-level annotations across 15 thematic categories (worship, ethics, legal rulings, family law, social justice, eschatology, stories of prophets, etc.) and 8 legal-domain sub-categories for verses with fiqh implications. All annotations reviewed by two scholars with ijaza (certification) in Quranic sciences.
  2. Query-passage relevance annotation: 12,000 query-passage pairs annotated for three relevance levels (directly answers, thematically related, not relevant) and one specialist dimension (Islamic scholarly consensus on the connection). Annotators held degrees in Islamic studies or Sharia law.

Post-deployment results after retraining on the annotated corpus:

The project required 22 weeks because of the scholar review layer — finding scholars with both Islamic science credentials and annotation protocol familiarity is the primary bottleneck in classical Arabic AI projects. The annotated corpus is now maintained with quarterly refresh cycles as new search query patterns emerge.

Annotator Qualification Requirements: What Actually Works

The annotator requirement for Classical Arabic NLP is not “native Arabic speaker.” It is not even “university-educated Arabic speaker.” The correct requirement is task-specific, and understanding the distinction prevents expensive annotation failures.

For morphological tagging of Classical Arabic texts: Annotators need formal training in classical Arabic grammar (nahw and sarf at the level of Al-Ajrumiyyah or higher). This is typically found in annotators with Islamic education backgrounds — not necessarily full Islamic science degrees, but exposure to classical Arabic grammar instruction. Egyptian and Saudi universities with strong Islamic studies faculties produce annotators with this background.

For NER on Hadith and historical texts: Annotators additionally need familiarity with classical Arabic naming conventions — the kunya (Abu X / Umm X form), the laqab (honorific), the nasab (genealogical chain). Without this, annotators systematically miss entity boundaries and misattribute names. A one-day calibration training on naming conventions is not sufficient — annotators need background knowledge to identify entities in context.

For Quranic thematic and legal annotation: Annotators need Islamic studies training at the level of a bachelor's degree in Sharia or equivalent. All annotations should be reviewed by at least one scholar with ijaza in the relevant Islamic science.

For Tajweed annotation for speech AI: Annotators need Tajweed certification — specifically, an ijaza in Tajweed recitation. This is a distinct qualification from Islamic science training and requires finding a different annotator pool.

Heritage Manuscripts: The Frontier of Classical Arabic NLP

Beyond the canonical Islamic texts, an enormous corpus of Classical Arabic manuscripts exists in libraries across the Arab world, Turkey, and Iran — the collections of the King Abdulaziz Public Library, the Süleymaniye Library in Istanbul, the Bodleian Library at Oxford, and hundreds of smaller archives. The King Abdulaziz Public Library alone holds over 300,000 manuscripts, and digitisation programmes funded by Saudi Vision 2030 are making these collections available.

Manuscript NLP annotation adds two challenges beyond standard Classical Arabic annotation. First, HTR (Handwritten Text Recognition) errors: Digitised manuscripts must be transcribed via HTR before NLP annotation begins. HTR models for Arabic manuscripts achieve 80–92% character accuracy but introduce systematic errors that annotators must understand to read correctly. Second, scribal variation: Pre-print manuscripts were copied by hand, and scribal traditions varied by region and period — spelling conventions, word boundaries, and abbreviations differ across the manuscript tradition.

Manuscript annotation projects in 2026 are typically structured as three-layer pipelines: HTR transcription → Classical Arabic annotator review and correction → domain-expert annotation for NLP tasks. Annotation throughput is significantly lower than for digital text — 200–400 tokens per annotator-hour for manuscript review, compared to 800–1,200 tokens per hour for standard Classical Arabic digital text.

What to Ask Before Commissioning Classical Arabic Annotation

When evaluating an Arabic NLP annotation partner for classical or Quranic Arabic projects, these questions surface genuine capability:

For broader Arabic NLP dataset building, see our post on sourcing and building Arabic NLP datasets and our guide to Arabic instruction-tuning data. Our Arabic NLP annotation service covers the full spectrum from modern dialect annotation through classical and Quranic text, with qualified annotator pools for each domain.

Frequently Asked Questions

What is Classical Arabic NLP and how does it differ from Modern Standard Arabic NLP?
Classical Arabic NLP applies NLP tasks to pre-modern texts: Quranic text, Hadith, classical jurisprudence, and heritage manuscripts. It differs from MSA NLP in three ways: a fuller case-marking system (i'rab); vocabulary and idioms absent from modern corpora; and fully diacritised text where correct reading can be theologically significant. Leading Arabic models achieve 45–62% F1 on Classical morphological tagging vs 85–92% on MSA tasks — a training data gap, not an architectural one.
What qualifications do annotators need for Quranic Arabic annotation?
Quranic annotation requires Tajweed proficiency for speech tasks (certified recitation knowledge) and classical Arabic grammar training (nahw and sarf) for text tasks. Thematic classification and legal ruling (ahkam) annotation additionally require Islamic science expertise at bachelor's degree level or higher, plus scholar review for theologically sensitive determinations. Crowdsourced or general Arabic-speaker annotation produces systematic errors on these tasks.
Can modern Arabic NLP models be fine-tuned for Classical Arabic tasks?
Yes, but with significant annotation investment. AraBERT and CAMeLBERT provide useful starting points — they have some Classical Arabic exposure through Quranic text in training corpora. Zero-shot performance is 45–62% F1; fine-tuning with 5,000–50,000 qualified-annotator tokens raises this to 80–92%, depending on task complexity. Morphological tagging converges with less data than NER on historical texts.
What NLP tasks require Classical or Quranic Arabic annotation?
Seven main tasks: morphological analysis (root extraction, POS tagging, case marking); NER on Hadith and historical texts; Quranic thematic and legal classification; speech synthesis Tajweed annotation; tafsir and commentary alignment; Hadith isnad extraction; and Sharia contract compliance analysis for Islamic fintech. Each task has distinct annotator qualification requirements.
What does Classical Arabic annotation cost compared to modern Arabic?
Classical Arabic annotation carries a 2–4× premium over modern Arabic. Morphological tagging by classically-trained linguists: AUD $0.25–0.55 per token. Islamic science expert annotation (Hadith classification, fiqh rulings): AUD $0.60–1.40 per item. Tajweed annotation for speech AI: AUD $0.35–0.75 per verse segment. The premium reflects annotator scarcity rather than task complexity.
Free Sample · 24-48 hours

Start Your Classical Arabic Annotation Project

Tell us your text domain (Quranic, Hadith, fiqh, manuscripts), required annotation tasks, and volume. We will match you with classically-trained Arabic annotators and provide a project plan within 48 hours.

No commitment. NDA available on request. We respond within 24 hours, often the same day for Gulf-region inquiries.

Neel Bennett

AI Annotation Specialist at AI Taggers

Neel has over 8 years of experience in AI training data and machine learning operations. He specializes in helping enterprises build high-quality datasets for computer vision and NLP applications across healthcare, automotive, and retail industries.

Connect on LinkedIn