Arabic & MENALLM Evaluation

The Open Arabic LLM Leaderboard (OALL) Explained for AI Teams

OALL is the Arabic LLM community's primary evaluation benchmark suite — broader than ArabicMMLU alone and the closest thing the Arabic AI ecosystem has to an independent quality standard. Here is what it measures, who leads on each dimension, and how to use it as an annotation diagnostic.

16 August 202612 min read

Quick answer

The Open Arabic LLM Leaderboard (OALL) is a multi-benchmark evaluation suite maintained by the Arabic NLP research community. It evaluates Arabic LLMs across ArabicMMLU (multitask understanding), ACVA (Arabic cultural values alignment), AlGhafa (diverse NLP tasks), and dialect comprehension tasks. Top 2026 performers include Jais-30B, AceGPT-70B, and frontier multilingual models, with Arabic-specialist models leading on cultural alignment and dialectal tasks. OALL scores directly map to training data gaps: low ACVA scores indicate insufficient authentic Arabic cultural content; low dialect scores indicate underrepresentation of native-speaker data for Gulf, Egyptian, or Levantine registers.

What Is OALL and Why Does the Arabic AI Community Need It?

The Open Arabic LLM Leaderboard (OALL) was created to address a specific gap: the absence of a standardised, multi-dimensional evaluation framework for Arabic language models that could function as an independent quality signal for the research community, AI companies, and government AI programmes across the Arab world.

Prior to OALL, evaluating Arabic LLMs required running multiple disparate benchmarks separately: ArabicMMLU for multitask understanding, ACVA for cultural alignment, AlGhafa for NLP capability, and various ad hoc dialect tests. OALL aggregates these into a single leaderboard with consistent evaluation methodology, making it possible to compare models across all dimensions simultaneously and track progress over time.

The leaderboard is maintained by researchers at multiple Arab universities and AI labs, with contributions from the Hugging Face Arabic community. It uses standardised evaluation protocols — same prompts, same few-shot settings, same scoring methodology — across all submitted models, making it the most reliable comparative signal currently available for Arabic LLM capability.

OALL's Benchmark Components: What Each One Measures

Understanding what each OALL component measures is essential for using it as an annotation strategy tool. A low score on one component points to a different type of training data gap than a low score on another.

1

ArabicMMLU — Multitask Understanding (Largest Component)

40 subject areas in Modern Standard Arabic and major dialects. Tests factual knowledge, reasoning, and language comprehension across STEM, humanities, law, and Islamic studies. Low ArabicMMLU scores on specific subjects indicate gaps in domain-specific Arabic training data. Low scores on dialect subsets indicate insufficient dialectal data. This is the highest-weight component in the OALL composite score.

2

ACVA — Arabic Cultural Values and Alignment

62 cultural topics covering Islamic religious practice, MENA family structures, gender norms, social etiquette, and regional identity. ACVA tests whether a model generates outputs that align with the values and expectations of Arabic-speaking communities — not just whether it generates grammatically correct Arabic. This is the benchmark where translated English training data fails most visibly: a model trained on translated English SFT data may give technically correct but culturally misaligned Arabic responses that would be unacceptable in a Gulf government or family health context.

3

AlGhafa — Diverse Arabic NLP Tasks

A multi-task benchmark covering Arabic named entity recognition (NER), sentiment analysis, machine translation quality evaluation, dialect identification, and question answering. AlGhafa tests NLP capability breadth. Models that score well on ArabicMMLU but poorly on AlGhafa typically have strong language understanding but weak NLP task performance — a common profile for models fine-tuned heavily on question-answering formats without diversity in task types.

4

EXAMS-Arabic — Academic Exam Questions

Arabic-language secondary and university exam questions from Egyptian, Saudi Arabian, and Jordanian curricula. Tests whether a model understands Arabic in the register and vocabulary used in formal education — relevant for EdTech applications and for assessing whether a model would perform well in contexts where Arabic-medium academic training is the primary use case.

Who Leads on OALL in 2026 — and Why

The OALL leaderboard in mid-2026 reveals a pattern consistent with the ArabicMMLU-specific leaderboard: no single model leads on all four components simultaneously.

Model categoryArabicMMLUACVA culturalAlGhafa NLPDialect tasks
Frontier multilingual (GPT-4-class)High (72–78%)Moderate (62–68%)High (70–76%)Moderate (58–64%)
Arabic-specialist (Jais-30B, AceGPT-70B)Moderate (64–72%)High (71–78%)Moderate (65–72%)High (66–74%)
Fine-tuned open-weight Arabic modelsVaries (58–70%)Varies (55–72%)Varies (60–70%)Varies (48–68%)

The ACVA column is the most instructive. Frontier multilingual models score 8–12 points lower on ACVA than on ArabicMMLU — their cultural alignment in Arabic is weaker than their factual Arabic comprehension. This is a direct consequence of training data: a model trained on web-crawled Arabic text plus translated English SFT data will produce culturally misaligned outputs even if it scores well on academic-style question answering. The fix is authentic, native-speaker-annotated Arabic instruction-tuning data that embeds Gulf and MENA cultural norms into model behaviour — exactly what our Arabic NLP annotation service is designed to produce.

Need Arabic annotation that improves OALL scores?

AI Taggers provides native-speaker Arabic annotation targeting ACVA cultural alignment, ArabicMMLU subject gaps, and dialectal comprehension. PDPL-compliant for KSA and GCC projects.

Explore Arabic NLP annotation

Case Study: UAE Conversational AI Team — ACVA Score from 59% to 74%

A Dubai-based conversational AI company building a customer-service assistant for the UAE healthcare sector discovered a critical OALL weakness during a pre-launch evaluation round. Their model — a fine-tuned variant of a strong open-weight base model — was performing acceptably on ArabicMMLU (67.4% overall) but failing badly on ACVA (59.1%), particularly on questions relating to Islamic medical ethics, patient communication norms in Gulf healthcare contexts, and gender-appropriate communication styles in a UAE hospital setting.

The issue was identifiable in hindsight: their fine-tuning data was almost entirely translated English healthcare dialogue. The English source data used clinical communication styles, directness levels, and gender interaction norms appropriate for Western healthcare — none of which transferred correctly to Emirati and broader Gulf Arabic healthcare contexts.

OALL scores before intervention

ArabicMMLU overall

67.4%

ACVA cultural alignment

59.1%

Khaleeji dialect tasks

53.7%

AlGhafa NLP suite

64.2%

AI Taggers designed a two-part annotation programme over ten weeks:

  1. 11,000 records of UAE healthcare dialogue SFT data — annotated by native Emirati and Gulf-Arabic annotators with familiarity with UAE healthcare communication norms. Tasks included patient greeting protocols in Gulf Arabic, culturally appropriate phrasing for sensitive health topics, Islamic ethical framing for clinical recommendations, and gender-appropriate register selection.
  2. 7,500 records of Khaleeji conversational instruction-tuning data — covering UAE and broader GCC registers for service sector dialogue, with special attention to indirect communication styles and formulaic politeness expressions absent from MSA corpora.

Post-intervention OALL results after fine-tuning on this dataset:

OALL scores after intervention

ArabicMMLU overall

71.8% (+4.4pp)

ACVA cultural alignment

74.3% (+15.2pp)

Khaleeji dialect tasks

69.4% (+15.7pp)

AlGhafa NLP suite

68.9% (+4.7pp)

The 15.2-point ACVA improvement was driven almost entirely by the culturally-annotated healthcare dialogue data. In production piloting, patient satisfaction scores for the AI assistant rose from 3.1/5 to 4.3/5 (measured by post-consultation survey) — with the most common free-text complaint in the pre-intervention period being that the assistant "spoke like a textbook" rather than like a respectful Gulf Arabic interlocutor. The cultural alignment data fixed this in a single fine-tuning round, without any architecture changes.

Using OALL as an Annotation Strategy Tool

For Arabic AI teams, OALL serves three practical functions beyond competitive benchmarking:

1. Pre-annotation diagnostic

Before commissioning any annotation, run your current model (or base model) against the full OALL suite. Generate a breakdown by benchmark component and by subject area within ArabicMMLU. This takes approximately four hours on a standard GPU and produces a ranked list of your model's weakest areas — which is your annotation priority list.

A model scoring 65% on ArabicMMLU but 55% on ACVA needs culturally-grounded Arabic SFT data before anything else. A model scoring 70% on ArabicMMLU but 48% on Khaleeji dialect tasks needs Gulf-native annotators producing dialectal dialogue data. These are different annotation products with different annotator requirements, and OALL tells you which one to prioritise.

2. Post-annotation quality gate

Re-run OALL after each annotation and fine-tuning round. Improvements in the targeted benchmark areas confirm that the annotation data is working. A round that improves ArabicMMLU but regresses ACVA signals that the SFT data was culturally neutral (MSA-dominant) when it needed to be culturally embedded. OALL's multi-benchmark structure catches these trade-offs that single-benchmark evaluation misses.

3. Comparative positioning for procurement

Gulf government AI procurement is increasingly requiring OALL benchmark submissions alongside vendor proposals. A model that can demonstrate 70%+ on ACVA alongside strong ArabicMMLU performance signals that it was trained on authentic Arabic data — not translated English content — which is a meaningful differentiator for sensitive government and healthcare applications.

For a deeper look at how Saudi Arabia's AI ambitions are driving demand for Arabic training data at scale, see our analysis of Vision 2030 and the sovereign Arabic LLM push. For teams building on the ArabicMMLU component specifically, our post on ArabicMMLU top models and what they reveal about training data covers the subject-area breakdown in detail. See also our Arabic data labeling service for end-to-end annotation pipelines targeting specific OALL benchmark gaps.

The OALL Data Gap the Leaderboard Reveals

Every model currently on the OALL leaderboard has the same weakness: ACVA scores trail ArabicMMLU scores by 8–15 points. This is a reliable signal that the Arabic AI training data ecosystem is undersupplied with authentic, native-speaker-annotated cultural data relative to MSA factual content.

According to analysis by the OALL maintainers, the primary cause is that the Arabic NLP data commons is built predominantly on MSA news corpora, legal texts, and translated English datasets — all of which are culturally thin relative to the rich oral and social culture of Gulf and MENA communities. Filling this gap requires native-speaker annotators with genuine cultural competence, not just language fluency. A fluent MSA speaker from an Egyptian academic background will not reliably annotate Emirati social norms or Saudi Islamic jurisprudence questions correctly without domain expertise.

This is why the ACVA gap persists even for frontier multilingual models with enormous total training budgets: scale cannot substitute for cultural authenticity when the underlying data is culturally thin. As we explain in our guide to Arabic text annotation tools and best practices, the people question — which annotators, with what native competence — matters far more than the platform or tooling question.

Frequently Asked Questions

What is the Open Arabic LLM Leaderboard (OALL)?
OALL is a community-maintained evaluation suite for Arabic LLMs that combines multiple benchmarks: ArabicMMLU (40-subject multitask understanding), ACVA (Arabic cultural values alignment), AlGhafa (diverse NLP tasks), and EXAMS-Arabic (academic exam questions). It provides a multi-dimensional view of Arabic model capability and is maintained by Arabic NLP researchers across the Arab world.
Which models top OALL in 2026?
No single model leads on all four OALL components. Frontier multilingual models (GPT-4-class) achieve the highest ArabicMMLU and AlGhafa scores. Arabic-specialist models including Jais-30B (G42, Abu Dhabi) and AceGPT-70B lead on ACVA cultural alignment and Khaleeji dialect tasks. For Gulf-facing AI products, the ACVA and dialect leaderboards are more relevant than aggregate rankings.
Why is ACVA the most important OALL benchmark for Gulf AI?
ACVA (Arabic Cultural Values and Alignment) tests whether a model generates outputs consistent with MENA cultural norms across 62 topics. A model that scores well on ArabicMMLU but poorly on ACVA will produce culturally misaligned responses in production — technically correct Arabic that violates Gulf social norms. This creates real product problems in healthcare, government, and financial services contexts where cultural appropriateness is non-negotiable.
How do I use OALL to decide what annotation to commission?
Run your model against the full OALL suite, generate a breakdown by benchmark and by subject area within ArabicMMLU. Your annotation priority is your lowest-scoring cluster. Low ACVA: commission culturally-grounded Arabic SFT data with native annotators who have relevant domain competence. Low Khaleeji dialect: commission Gulf-native conversational annotation. Low ArabicMMLU on specific subjects: commission domain-specific Arabic training data in those subject areas.
Is OALL the same as the Hugging Face Arabic leaderboard?
OALL is inspired by the Hugging Face Open LLM Leaderboard methodology but is a separate initiative maintained by the Arabic NLP community. It uses benchmarks specific to Arabic, including ACVA which has no English equivalent. Some models submitted to OALL also appear in Hugging Face leaderboards, but the evaluation criteria and benchmark components are different.
Does PDPL compliance affect which annotation vendor I can use for OALL-targeting data?
Yes, if your training data includes personal data of KSA residents. Saudi Arabia's PDPL (Personal Data Protection Law) governs how annotation vendors may handle, store, and transfer personal Arabic data. For KSA-facing AI products, annotation data sourced from Saudi users, documents, or conversations must be handled under PDPL-compliant workflows with SDAIA-aligned data residency controls. AI Taggers provides PDPL-compliant annotation for all KSA and Gulf AI projects.
Free Sample · 24-48 hours

Get a quote for OALL-targeted Arabic annotation

Tell us your OALL benchmark gaps, target dialect, and volume. We'll respond with a scoped proposal within one business day.

No commitment. NDA available on request. We respond within 24 hours, often the same day for Gulf-region inquiries.

Neel Bennett

AI Annotation Specialist at AI Taggers

Neel has over 8 years of experience in AI training data and machine learning operations. He specializes in helping enterprises build high-quality datasets for computer vision and NLP applications across healthcare, automotive, and retail industries.

Connect on LinkedIn