Arabic & MENALLM Benchmarking

ArabicMMLU Leaderboard: Which Model Tops Arabic Benchmarks (and Why It Matters for Your Data)

The ArabicMMLU leaderboard exposes exactly where Arabic LLMs succeed and fail across 40 subject areas and multiple dialects. Understanding which models lead — and why they lead — tells you precisely what training data your own Arabic AI project is missing.

16 August 202613 min read

Quick answer

ArabicMMLU is a 40-subject multitask benchmark for Arabic language understanding, covering Modern Standard Arabic and major dialects. As of mid-2026, frontier multilingual models (GPT-4-class, Gemini Ultra-class) score 72–78% overall, while Arabic-specialist models such as Jais-30B and AceGPT-70B lead on dialectal and domain-specific subsets. The critical finding: all models drop 15–25 percentage points from MSA to dialectal tasks — a gap driven by training data, not architecture, and correctable through targeted native-speaker annotation.

What Is ArabicMMLU and Why Does It Matter?

ArabicMMLU (Arabic Massive Multitask Language Understanding) is modelled on the widely used English MMLU benchmark, adapted for the unique challenges of Arabic: its diglossic register system (Modern Standard Arabic versus spoken dialects), its right-to-left script, its morphological richness, and the significant cultural and religious domain knowledge required to answer questions correctly.

The benchmark covers 40 subject areas including Islamic jurisprudence, KSA and UAE law, Arabic literature, STEM disciplines in Arabic, history of the Arab world, and colloquial dialect comprehension tasks. Questions are multiple-choice with four options. Models receive a single aggregate score plus a subject-by-subject breakdown, which is where the actionable insight lives.

According to the ArabicMMLU paper (Koto et al., 2024), even the best models at release time scored significantly below human performance on Arabic-cultural and dialectal subsets. Human performance on the cultural and Islamic jurisprudence categories is approximately 90–95%; model performance at the time of publication ranged from 48–67% on those same categories — a gap that reveals precisely where Arabic AI training data is thin.

Current Leaderboard: Who Leads and on Which Tasks

As of mid-2026, the ArabicMMLU leaderboard reflects a two-tier structure:

Tier 1: Frontier multilingual models (72–78% aggregate)

GPT-4-class and Gemini Ultra-class systems achieve the highest aggregate ArabicMMLU scores. Their strength is breadth — they perform consistently across MSA STEM, humanities, and law subjects. Their weakness is Khaleeji dialect tasks and culturally specific Islamic jurisprudence questions, where aggregate scores drop to 58–64%. These models achieve high overall scores by performing very well on the larger, MSA-heavy subject areas that dominate the aggregate.

Tier 2: Arabic-specialist models (64–72% aggregate, stronger on dialect subsets)

Models including Jais-30B (developed by G42/Inception in Abu Dhabi), AceGPT-70B, and Silma-9B score 5–10 points below frontier multilingual models on aggregate but outperform them on Khaleeji, Islamic law, and Gulf regulatory tasks by 6–14 points. For teams building KSA-facing products — government services, legal AI, fintech compliance — these models provide better out-of-the-box performance in the domains that matter, despite lower aggregate rankings.

A finding consistent across the leaderboard: every model drops 15–25 percentage points from MSA task performance to dialectal task performance. This is not a model capability problem. It is a training data problem. Models trained predominantly on MSA web text, news corpora, and translated English datasets perform well on MSA questions because that is where their training data is concentrated. Dialectal Arabic — Gulf Khaleeji, Egyptian Aamiyya, Levantine, Iraqi, Maghrebi Darija — is underrepresented in every standard Arabic pre-training corpus.

Reading the Subject-Area Breakdown: Where Your Model Is Missing Data

The aggregate ArabicMMLU score is nearly useless for making annotation decisions. The subject-area and dialect breakdowns are where the insight is. Teams building Arabic AI for specific verticals should run their model against the relevant ArabicMMLU subsets and read the breakdown as a training data diagnostic report.

Common patterns seen in subject-area breakdowns:

For teams working with our Arabic NLP annotation services, a pre-engagement ArabicMMLU diagnostic run is the fastest way to identify where to allocate annotation budget. Each low-scoring subject cluster maps directly to an annotation category: dialectal data collection, domain-specific document annotation, or instruction-tuning data in the target register.

Need Arabic annotation that targets your benchmark gaps?

AI Taggers provides native-speaker Arabic annotation across all dialects — Khaleeji, Egyptian, Levantine, MSA, and Maghrebi — with domain coverage for law, Islamic studies, STEM, and regulatory documents. PDPL-compliant for KSA projects.

Explore Arabic NLP annotation

Case Study: KSA Legal AI Startup — From 58% to 71% on ArabicMMLU (Islamic Jurisprudence Subset)

A Riyadh-based legal-tech startup building an AI assistant for Sharia-compliant contract review approached AI Taggers after their fine-tuned Arabic LLM plateaued on benchmark performance. Their model was based on a strong open-weight Arabic foundation model and had been fine-tuned on translated English legal data — a common starting point that creates predictable ArabicMMLU weaknesses.

Pre-intervention ArabicMMLU scores

Overall aggregate

58.3%

Islamic jurisprudence subset

44.7%

Khaleeji dialect tasks

41.2%

KSA regulatory law

52.8%

The diagnostic was clear: the model was performing below chance level on Islamic jurisprudence questions (chance for 4-option multiple choice is 25%, but the model was at 44.7%, indicating some learning but severe gaps in classical Arabic legal reasoning). The Khaleeji task score of 41.2% reflected near-zero exposure to Gulf-register Arabic during fine-tuning.

AI Taggers commissioned two annotation workstreams over eight weeks:

  1. 14,000 records of KSA legal and Islamic jurisprudence SFT data — native-speaker annotators with dual competence in classical Arabic and Saudi legal practice. Tasks included Quranic terminology contextualisation, fiqh ruling classification, and contract-clause reading comprehension in formal Hejazi Arabic.
  2. 9,500 records of Khaleeji-register instruction-tuning data — Gulf-native annotators covering KSA, UAE, Kuwait, and Bahrain registers. Tasks included business-register question-answering, regulatory dialogue, and document comprehension in mixed formal/colloquial Gulf Arabic.

After fine-tuning on this targeted dataset (no changes to model architecture or pre-training):

Post-intervention ArabicMMLU scores

Overall aggregate

66.1% (+7.8pp)

Islamic jurisprudence subset

71.2% (+26.5pp)

Khaleeji dialect tasks

63.8% (+22.6pp)

KSA regulatory law

68.4% (+15.6pp)

The aggregate gain of 7.8 percentage points understates the improvement: the model moved from below-chance performance on its primary product use case to a competitive score in its target domain. In production, legal document classification accuracy on KSA court documents improved from 67.3% to 84.1%, reducing manual review workload by approximately 40% for the startup's pilot law firm clients.

What the Leaderboard Tells You About Your Annotation Priorities

Teams building Arabic AI for production use should read ArabicMMLU leaderboard standings less as a model selection guide and more as a map of where the Arabic AI training data ecosystem is currently weak. If even the best models score 58–64% on Khaleeji tasks, that means native Khaleeji annotation at scale is rare and valuable. Whoever builds it gains a durable advantage that cannot be replicated by scaling a model trained on MSA text.

According to research from the ArabicMMLU and OALL teams, the domains with the largest gap between human performance and model performance in 2024–2025 were: Islamic jurisprudence (human ~92%, best model ~62%), KSA regulatory law (human ~88%, best model ~58%), and Khaleeji colloquial comprehension (human ~91%, best model ~56%). These gaps persist in 2026 for any model not specifically trained on native-annotated dialectal data.

The practical implication for MENA AI teams: if your Arabic LLM product targets KSA consumers, Gulf government clients, or Islamic finance institutions, your ArabicMMLU scores in those subject areas are your product quality signal. Closing the gap requires Arabic NLP annotation from credentialed native speakers with domain competence — not more translated English data or synthetic Arabic generation from MSA base models.

Arabic-Specialist Models vs Frontier Multilingual: Which Should You Fine-Tune?

Teams often ask whether to start from a frontier multilingual model (higher aggregate ArabicMMLU, wider world knowledge) or an Arabic-specialist model (better Khaleeji and domain-specific baseline, smaller context window in some cases).

The answer depends on your product's dialect and domain mix:

Regardless of base model choice, the annotation strategy is the same: run an ArabicMMLU diagnostic, identify the low-scoring subject clusters, and commission native-speaker annotation in those exact areas. As our guide to building custom Arabic benchmarks explains, teams that supplement ArabicMMLU with product-specific evaluation sets get a much cleaner signal about what annotation to prioritise.

Vision 2030 and the Sovereign Arabic LLM Push

The ArabicMMLU leaderboard is not just a research artefact — it is a competitive battleground for Vision 2030. Saudi Arabia's SDAIA, Aramco Digital, and PIF-backed AI initiatives have all framed sovereign Arabic LLM capability as a national priority. The models competing on ArabicMMLU are, in many cases, competing for government AI contracts that depend on demonstrated benchmark performance in precisely the Islamic jurisprudence, KSA law, and Khaleeji language categories where current models are weakest.

As we covered in detail in our analysis of Saudi Arabia's sovereign LLM push, the annotation market this creates is substantial. Teams that build high-quality Arabic training datasets — particularly in the benchmark-weak categories — are positioned to serve both the KSA government AI programmes and the commercial Arabic LLM companies competing for Gulf market share.

For practical guidance on sourcing and building Arabic training datasets from scratch, our post on Arabic NLP dataset sourcing covers licensing, dialect balance, and provenance requirements in depth. See also our Arabic data labeling service page for production-scale engagement options.

Frequently Asked Questions

What is ArabicMMLU?
ArabicMMLU is a massively multitask language understanding benchmark for Arabic covering 40 subject areas in Modern Standard Arabic and major dialects including Gulf Khaleeji, Egyptian, and Levantine. Models are evaluated with multiple-choice questions and scored by subject area, revealing domain and dialect gaps in Arabic LLM training data.
Which model tops the ArabicMMLU leaderboard in 2026?
Frontier multilingual models (GPT-4-class, Gemini Ultra-class) achieve the highest aggregate scores at 72–78%. Arabic-specialist models including Jais-30B and AceGPT-70B lead on Khaleeji dialect and Islamic jurisprudence subsets despite lower aggregate rankings. The leaderboard is most useful when read by subject area, not aggregate.
Why do Arabic LLMs score so much lower on dialect tasks than MSA tasks?
Because virtually all Arabic pre-training corpora are dominated by MSA news text, Wikipedia, and translated English data. Dialectal Arabic — Gulf Khaleeji, Egyptian, Levantine, Iraqi, Maghrebi — is severely underrepresented. The 15–25 point drop from MSA to dialectal tasks seen across the leaderboard is a training data gap, not a model architecture gap.
How can I use ArabicMMLU to prioritise my annotation budget?
Run your model against ArabicMMLU and generate a subject-area and dialect breakdown. Identify the lowest-scoring clusters — these are your training data gaps. Commission native-speaker annotation in those exact domains. The Khaleeji, Islamic jurisprudence, and KSA regulatory law categories are consistently the weakest across all publicly available models and benefit most from targeted native-speaker annotation.
What is PDPL and does it affect Arabic annotation?
PDPL (Personal Data Protection Law) is Saudi Arabia's data privacy law administered by SDAIA. If your annotation data includes KSA residents' personal information — names, locations, financial records — your annotation vendor must handle it under PDPL-compliant data residency and access controls. AI Taggers provides PDPL-compliant annotation workflows for all KSA AI projects.
Is ArabicMMLU the same as the OALL leaderboard?
No. The Open Arabic LLM Leaderboard (OALL) is a broader evaluation suite that includes ArabicMMLU plus additional benchmarks such as ACVA, AlGhafa, and dialect-specific tasks. ArabicMMLU is one component within OALL. Models can rank differently on the two leaderboards depending on their training data distribution across MSA versus dialect coverage.
Free Sample · 24-48 hours

Get a quote for Arabic LLM annotation

Tell us your target dialect, ArabicMMLU weak spots, and volume. We'll respond with a scoped proposal within one business day.

No commitment. NDA available on request. We respond within 24 hours, often the same day for Gulf-region inquiries.

Neel Bennett

AI Annotation Specialist at AI Taggers

Neel has over 8 years of experience in AI training data and machine learning operations. He specializes in helping enterprises build high-quality datasets for computer vision and NLP applications across healthcare, automotive, and retail industries.

Connect on LinkedIn