Quick answer
ArabicMMLU is a 40-subject multitask benchmark for Arabic language understanding, covering Modern Standard Arabic and major dialects. As of mid-2026, frontier multilingual models (GPT-4-class, Gemini Ultra-class) score 72–78% overall, while Arabic-specialist models such as Jais-30B and AceGPT-70B lead on dialectal and domain-specific subsets. The critical finding: all models drop 15–25 percentage points from MSA to dialectal tasks — a gap driven by training data, not architecture, and correctable through targeted native-speaker annotation.
What Is ArabicMMLU and Why Does It Matter?
ArabicMMLU (Arabic Massive Multitask Language Understanding) is modelled on the widely used English MMLU benchmark, adapted for the unique challenges of Arabic: its diglossic register system (Modern Standard Arabic versus spoken dialects), its right-to-left script, its morphological richness, and the significant cultural and religious domain knowledge required to answer questions correctly.
The benchmark covers 40 subject areas including Islamic jurisprudence, KSA and UAE law, Arabic literature, STEM disciplines in Arabic, history of the Arab world, and colloquial dialect comprehension tasks. Questions are multiple-choice with four options. Models receive a single aggregate score plus a subject-by-subject breakdown, which is where the actionable insight lives.
According to the ArabicMMLU paper (Koto et al., 2024), even the best models at release time scored significantly below human performance on Arabic-cultural and dialectal subsets. Human performance on the cultural and Islamic jurisprudence categories is approximately 90–95%; model performance at the time of publication ranged from 48–67% on those same categories — a gap that reveals precisely where Arabic AI training data is thin.
Current Leaderboard: Who Leads and on Which Tasks
As of mid-2026, the ArabicMMLU leaderboard reflects a two-tier structure:
Tier 1: Frontier multilingual models (72–78% aggregate)
GPT-4-class and Gemini Ultra-class systems achieve the highest aggregate ArabicMMLU scores. Their strength is breadth — they perform consistently across MSA STEM, humanities, and law subjects. Their weakness is Khaleeji dialect tasks and culturally specific Islamic jurisprudence questions, where aggregate scores drop to 58–64%. These models achieve high overall scores by performing very well on the larger, MSA-heavy subject areas that dominate the aggregate.
Tier 2: Arabic-specialist models (64–72% aggregate, stronger on dialect subsets)
Models including Jais-30B (developed by G42/Inception in Abu Dhabi), AceGPT-70B, and Silma-9B score 5–10 points below frontier multilingual models on aggregate but outperform them on Khaleeji, Islamic law, and Gulf regulatory tasks by 6–14 points. For teams building KSA-facing products — government services, legal AI, fintech compliance — these models provide better out-of-the-box performance in the domains that matter, despite lower aggregate rankings.
A finding consistent across the leaderboard: every model drops 15–25 percentage points from MSA task performance to dialectal task performance. This is not a model capability problem. It is a training data problem. Models trained predominantly on MSA web text, news corpora, and translated English datasets perform well on MSA questions because that is where their training data is concentrated. Dialectal Arabic — Gulf Khaleeji, Egyptian Aamiyya, Levantine, Iraqi, Maghrebi Darija — is underrepresented in every standard Arabic pre-training corpus.
Reading the Subject-Area Breakdown: Where Your Model Is Missing Data
The aggregate ArabicMMLU score is nearly useless for making annotation decisions. The subject-area and dialect breakdowns are where the insight is. Teams building Arabic AI for specific verticals should run their model against the relevant ArabicMMLU subsets and read the breakdown as a training data diagnostic report.
Common patterns seen in subject-area breakdowns:
- Islamic jurisprudence (fiqh) weak: indicates a gap in Arabic religious text annotation. The model lacks exposure to classical Arabic legal reasoning, Quranic terminology, and hadith citation structures.
- Khaleeji dialect tasks weak (below 55%): indicates MSA-dominated pre-training with minimal native Gulf Arabic data. The most common gap for Saudi-facing products.
- KSA/UAE regulatory law weak: indicates no local regulatory document annotation. Kingdom-level laws, SAMA guidelines, and NDMO frameworks use formal Arabic registers and terminology absent from general web corpora.
- Egyptian colloquial tasks weak: indicates lack of Egyptian Arabic social media, customer-service transcripts, or conversational data annotation.
- STEM in Arabic moderate (60–68%): indicates reasonable MSA coverage but gaps in Arabic-medium education text from Gulf universities, where scientific terminology differs from Egyptian or Levantine conventions.
For teams working with our Arabic NLP annotation services, a pre-engagement ArabicMMLU diagnostic run is the fastest way to identify where to allocate annotation budget. Each low-scoring subject cluster maps directly to an annotation category: dialectal data collection, domain-specific document annotation, or instruction-tuning data in the target register.
Need Arabic annotation that targets your benchmark gaps?
AI Taggers provides native-speaker Arabic annotation across all dialects — Khaleeji, Egyptian, Levantine, MSA, and Maghrebi — with domain coverage for law, Islamic studies, STEM, and regulatory documents. PDPL-compliant for KSA projects.
Explore Arabic NLP annotationCase Study: KSA Legal AI Startup — From 58% to 71% on ArabicMMLU (Islamic Jurisprudence Subset)
A Riyadh-based legal-tech startup building an AI assistant for Sharia-compliant contract review approached AI Taggers after their fine-tuned Arabic LLM plateaued on benchmark performance. Their model was based on a strong open-weight Arabic foundation model and had been fine-tuned on translated English legal data — a common starting point that creates predictable ArabicMMLU weaknesses.
Pre-intervention ArabicMMLU scores
Overall aggregate
58.3%
Islamic jurisprudence subset
44.7%
Khaleeji dialect tasks
41.2%
KSA regulatory law
52.8%
The diagnostic was clear: the model was performing below chance level on Islamic jurisprudence questions (chance for 4-option multiple choice is 25%, but the model was at 44.7%, indicating some learning but severe gaps in classical Arabic legal reasoning). The Khaleeji task score of 41.2% reflected near-zero exposure to Gulf-register Arabic during fine-tuning.
AI Taggers commissioned two annotation workstreams over eight weeks:
- 14,000 records of KSA legal and Islamic jurisprudence SFT data — native-speaker annotators with dual competence in classical Arabic and Saudi legal practice. Tasks included Quranic terminology contextualisation, fiqh ruling classification, and contract-clause reading comprehension in formal Hejazi Arabic.
- 9,500 records of Khaleeji-register instruction-tuning data — Gulf-native annotators covering KSA, UAE, Kuwait, and Bahrain registers. Tasks included business-register question-answering, regulatory dialogue, and document comprehension in mixed formal/colloquial Gulf Arabic.
After fine-tuning on this targeted dataset (no changes to model architecture or pre-training):
Post-intervention ArabicMMLU scores
Overall aggregate
66.1% (+7.8pp)
Islamic jurisprudence subset
71.2% (+26.5pp)
Khaleeji dialect tasks
63.8% (+22.6pp)
KSA regulatory law
68.4% (+15.6pp)
The aggregate gain of 7.8 percentage points understates the improvement: the model moved from below-chance performance on its primary product use case to a competitive score in its target domain. In production, legal document classification accuracy on KSA court documents improved from 67.3% to 84.1%, reducing manual review workload by approximately 40% for the startup's pilot law firm clients.
What the Leaderboard Tells You About Your Annotation Priorities
Teams building Arabic AI for production use should read ArabicMMLU leaderboard standings less as a model selection guide and more as a map of where the Arabic AI training data ecosystem is currently weak. If even the best models score 58–64% on Khaleeji tasks, that means native Khaleeji annotation at scale is rare and valuable. Whoever builds it gains a durable advantage that cannot be replicated by scaling a model trained on MSA text.
According to research from the ArabicMMLU and OALL teams, the domains with the largest gap between human performance and model performance in 2024–2025 were: Islamic jurisprudence (human ~92%, best model ~62%), KSA regulatory law (human ~88%, best model ~58%), and Khaleeji colloquial comprehension (human ~91%, best model ~56%). These gaps persist in 2026 for any model not specifically trained on native-annotated dialectal data.
The practical implication for MENA AI teams: if your Arabic LLM product targets KSA consumers, Gulf government clients, or Islamic finance institutions, your ArabicMMLU scores in those subject areas are your product quality signal. Closing the gap requires Arabic NLP annotation from credentialed native speakers with domain competence — not more translated English data or synthetic Arabic generation from MSA base models.
Arabic-Specialist Models vs Frontier Multilingual: Which Should You Fine-Tune?
Teams often ask whether to start from a frontier multilingual model (higher aggregate ArabicMMLU, wider world knowledge) or an Arabic-specialist model (better Khaleeji and domain-specific baseline, smaller context window in some cases).
The answer depends on your product's dialect and domain mix:
- KSA-facing products (government, legal, Islamic finance): Arabic-specialist models such as Jais-30B provide a stronger dialectal and domain baseline to fine-tune from, reducing the annotation volume needed to reach production-quality performance on Khaleeji tasks.
- Multi-region MENA products: Frontier multilingual models maintain stronger MSA and cross-domain performance that supports broader reach, but require more targeted dialectal annotation to close specific regional gaps.
- Enterprise Arabic chatbots targeting mixed audiences: A hybrid approach — frontier model with targeted dialectal fine-tuning data — typically produces the best combined performance, given that most Gulf business users code-switch between formal MSA and Khaleeji register within a single conversation.
Regardless of base model choice, the annotation strategy is the same: run an ArabicMMLU diagnostic, identify the low-scoring subject clusters, and commission native-speaker annotation in those exact areas. As our guide to building custom Arabic benchmarks explains, teams that supplement ArabicMMLU with product-specific evaluation sets get a much cleaner signal about what annotation to prioritise.
Vision 2030 and the Sovereign Arabic LLM Push
The ArabicMMLU leaderboard is not just a research artefact — it is a competitive battleground for Vision 2030. Saudi Arabia's SDAIA, Aramco Digital, and PIF-backed AI initiatives have all framed sovereign Arabic LLM capability as a national priority. The models competing on ArabicMMLU are, in many cases, competing for government AI contracts that depend on demonstrated benchmark performance in precisely the Islamic jurisprudence, KSA law, and Khaleeji language categories where current models are weakest.
As we covered in detail in our analysis of Saudi Arabia's sovereign LLM push, the annotation market this creates is substantial. Teams that build high-quality Arabic training datasets — particularly in the benchmark-weak categories — are positioned to serve both the KSA government AI programmes and the commercial Arabic LLM companies competing for Gulf market share.
For practical guidance on sourcing and building Arabic training datasets from scratch, our post on Arabic NLP dataset sourcing covers licensing, dialect balance, and provenance requirements in depth. See also our Arabic data labeling service page for production-scale engagement options.
Frequently Asked Questions
What is ArabicMMLU?▼
Which model tops the ArabicMMLU leaderboard in 2026?▼
Why do Arabic LLMs score so much lower on dialect tasks than MSA tasks?▼
How can I use ArabicMMLU to prioritise my annotation budget?▼
What is PDPL and does it affect Arabic annotation?▼
Is ArabicMMLU the same as the OALL leaderboard?▼
Get a quote for Arabic LLM annotation
Tell us your target dialect, ArabicMMLU weak spots, and volume. We'll respond with a scoped proposal within one business day.
Neel Bennett
AI Annotation Specialist at AI Taggers
Neel has over 8 years of experience in AI training data and machine learning operations. He specializes in helping enterprises build high-quality datasets for computer vision and NLP applications across healthcare, automotive, and retail industries.
Connect on LinkedIn