Arabic & MENA

Arabic LLM Training Data RFP: Template and Vendor Scoring Sheet

Procuring Arabic LLM training data without a structured RFP is how teams end up with undialected MSA data delivered as "Gulf Arabic". This template gives you the right questions before you receive the first vendor pitch.

10 October 2026 14 min read

Direct Answer

Arabic LLM training data is the annotated instruction pairs, preference comparisons, and evaluation benchmarks used to fine-tune and align large language models for Arabic tasks. An RFP for this data should specify dialect mix, data type (SFT, RLHF, eval), IAA thresholds, annotator qualification requirements, PDPL compliance, and a mandatory pilot before volume commitment. Vendors that cannot specify their dialect pool composition should not advance past the first round.

Why Arabic LLM Data Procurement Needs a Formal RFP

Arabic is spoken by over 400 million people across 22 countries, yet it remains one of the most under-resourced major languages in existing LLM training corpora. Published analyses have found that Arabic accounts for less than 1% of content in the training sets of many major foundation models, despite Arabic being the fifth most spoken language globally (Ethnologue, 2024). This creates a significant opportunity — and a significant procurement risk.

The risk is dialect substitution: vendors who claim Arabic coverage but deliver data weighted toward Modern Standard Arabic (MSA), when Gulf-dialect or Egyptian-dialect coverage was what you paid for. Without an RFP that specifies dialect mix, IAA requirements, and annotator credentials, you have no contractual basis to dispute delivery when the data turns out to be commercially useless for your target market.

Saudi Arabia's Vision 2030 AI ambitions — including SDAIA's Arabic language initiatives and the sovereign LLM programmes funded through PIF-backed entities — have accelerated demand for Khaleeji-native SFT and RLHF data specifically. Khaleeji data is harder to source and commands a premium; vendors with genuine pools charge more, but vendors without them will still pitch for the work. An RFP is how you tell the difference.

Our Arabic LLM training data service covers SFT, RLHF, and evaluation data across Khaleeji, Egyptian, Levantine, and MSA dialects. The template below is drawn from the procurement structures that consistently surface the most qualified vendors in this space.

The RFP Template: Section by Section

Use these sections as the skeleton of your RFP document. Each section includes the specific questions that separate capable vendors from those who cannot deliver what they pitch.

01

Project Overview and Objectives

  • Purpose of the data (pre-training, SFT, RLHF, safety evaluation, domain adaptation)
  • Model architecture context (decoder-only, encoder-decoder, target context length)
  • Target use case and product domain (customer service, search, code, medical, legal)
  • Intended deployment geographies (KSA only, GCC-wide, pan-Arab, global with Arabic support)
  • Confidentiality requirements and any public-release restrictions
02

Data Type and Volume Requirements

  • SFT pairs: volume, average prompt length, average response length, domain split
  • RLHF preference comparisons: volume, K (number of responses ranked per prompt), domain
  • Evaluation items: volume, task types (MMLU-style, open generation, safety probes), gold-label methodology
  • Red-teaming examples: harm categories required (violence, misinformation, jailbreak, cultural offence)
  • Data format: JSONL schema, field names, encoding (UTF-8 with BOM or without)
03

Dialect and Language Requirements

  • Required dialect mix with percentage targets (e.g., 35% Gulf/Khaleeji, 30% Egyptian, 20% Levantine, 15% MSA formal)
  • Code-switching requirements (Arabic-English, Arabic-French for Maghrebi)
  • Register requirements (formal MSA, semi-formal, colloquial, social media)
  • Transliteration handling (Arabic-chat alphabet / Arabizi) if applicable
  • Dialect identification metadata: should each record carry a dialect label?
04

Quality and IAA Standards

  • Minimum inter-annotator agreement threshold by task type (specify kappa or Krippendorff's alpha)
  • Gold set construction methodology and ongoing annotator calibration process
  • Adjudication process for disagreements
  • Acceptable error rate per record type at delivery
  • QA reporting: what metrics and format are included in each delivery batch?
05

Annotator Qualifications

  • Dialect provenance verification method (self-report only is not acceptable)
  • Minimum educational or professional background for domain-specific tasks (medical, legal, financial)
  • NDA and confidentiality requirements for annotators
  • Device and security controls (managed devices, no-public-tool policies)
  • Experience with LLM annotation tasks specifically (SFT writing, preference ranking)
06

Compliance and Data Governance

  • PDPL compliance: DPA template, cross-border transfer mechanism, deletion policy
  • ISO 27001 certification or equivalent — certificate number and expiry
  • Data residency: where is annotation work performed physically?
  • Breach notification process and SLA
  • Subprocessor disclosure: do they use subcontractors for any portion of the work?
07

Pilot Design

  • Volume: 50–100 SFT pairs or 50 preference comparisons before volume commitment
  • Deliverables: labelled data + IAA report + annotator profile + guideline retrospective
  • Timeline: maximum 5 business days for a text pilot
  • Evaluation criteria: how will you assess the pilot output before proceeding?

Vendor Scoring Sheet

Score each vendor 1–5 on each criterion, multiply by the weight, and sum for a total out of 100. Run the pilot (Section 07 of the RFP) before finalising scores — pilot quality should be observable, not self-reported.

CriterionWeightWhat to Evaluate
Dialect coverage proof20%Annotator pool breakdown with minimum counts per dialect
IAA reporting capability20%Per-task IAA scores from recent comparable projects
Annotator credentials15%Vetting methodology, qualification test design, calibration process
PDPL compliance15%Signed DPA, cross-border transfer mechanism, deletion evidence
Pilot quality15%Quality of the 50–100-record pilot output (run this before scoring)
Guideline capability10%Sample annotation guidelines for a comparable Arabic LLM task
Pricing and flexibility5%Per-unit pricing, batch minimums, change-order process

Note: the "pricing and flexibility" weight is intentionally low — vendors with the strongest dialect coverage and IAA track record should not be ruled out on price alone in this market.

Case Study: GCC Telecom Regional Assistant

Project Snapshot

40,000
SFT Pairs
Khaleeji + MSA
Dialects
Telecom CS
Domain
Perplexity −43%
Outcome

A GCC-based telecommunications operator building a regional customer assistant needed 40,000 SFT pairs and 10,000 RLHF preference comparisons across Khaleeji and MSA Arabic. Without a formal RFP process, they engaged a vendor based on a pitch deck alone. The delivered data had no dialect labelling metadata, 34% of what was contracted as “Gulf Arabic” was later identified as Najdi/Hijazi MSA, and no IAA scores were provided with the delivery.

After a re-procurement using the RFP template structure above — specifying dialect mix, mandatory IAA reporting, and a 100-pair pilot before commitment — they received properly segmented SFT pairs with IAA ≥ 0.83 across all task types, full dialect provenance metadata per record, and PDPL-compliant data handling documentation. After retraining on the new data, perplexity on Khaleeji customer service test sets dropped from 38.4 to 21.7, and intent accuracy on held-out Gulf-dialect queries improved by 24 points.

The original procurement failure cost approximately four months and the full sunk cost of the first vendor engagement. The RFP structure added two weeks to the procurement timeline and eliminated the ambiguity that caused the failure.

Request an Arabic LLM Data Pilot

Before you issue an RFP, see what production-quality Arabic LLM annotation looks like. We will produce 50–100 SFT pairs or preference comparisons with full IAA reporting and dialect metadata — free, within 5 business days.

Request a Free Arabic LLM Data Pilot

What a Strong Vendor Response Looks Like

Strong responses to an Arabic LLM training data RFP share several characteristics that are easy to differentiate from weak ones:

Dialect pool specificity

STRONG

Provides a breakdown by dialect with annotator counts (e.g., 47 Khaleeji-native, 38 Egyptian-native, 29 Levantine-native) and describes their vetting methodology

WEAK

Claims 'Arabic native speakers' without dialect breakdown, or provides a single aggregate count

IAA evidence

STRONG

Attaches IAA scores from a comparable recent project (SFT writing: κ = 0.82; preference ranking: κ = 0.78)

WEAK

Claims high quality without data, or cites only generic internal QA processes

Sample guidelines

STRONG

Provides a sample annotation guideline document for a comparable Arabic LLM task, showing edge cases and dialect-specific decision rules

WEAK

Describes guidelines verbally but cannot produce a sample document

PDPL compliance

STRONG

Provides a DPA template for review, names the cross-border transfer mechanism used, and describes the deletion process with timelines

WEAK

States they are 'compliant with data privacy requirements' without specifics

Pilot offer

STRONG

Offers to run the pilot before contract signing, specifying deliverables (labelled data + IAA report + annotator profile)

WEAK

Requires a signed agreement before any pilot data is provided

For context on what Arabic LLM evaluation should look like after you have the data, see our guide to Arabic LLM evaluation benchmarks. For vendor selection in Arabic annotation more broadly, see our 10-check guide to choosing an Arabic annotation company. Our Arabic data labelling service page covers the full workflow from task brief to production delivery.

The Most Common Arabic LLM Data RFP Mistakes

Even teams who use an RFP often make structural errors that leave them exposed to the same dialect-quality risks. The most common:

Specifying 'Arabic' without a dialect split — every vendor can claim coverage
Not requiring a pilot before contract signing — you have no quality evidence until delivery
Accepting IAA figures without specifying the metric (kappa and raw agreement are not comparable)
Omitting PDPL requirements until after contract negotiation begins, when it becomes a re-scoping exercise
Scoring only on price — the cheapest Arabic LLM data vendor is almost always using MSA annotators for dialect tasks
Not specifying delivery format and schema — receiving JSONL with Arabic in visual order rather than logical order is a common and expensive surprise

Frequently Asked Questions

What is Arabic LLM training data?

Arabic LLM training data is the annotated instruction pairs, preference comparisons, and evaluation benchmarks used to fine-tune and align large language models for Arabic-language tasks. Quality data requires native-speaker annotators by dialect — not bilingual generalists — and IAA measurement at each task type.

What should an RFP for Arabic LLM training data include?

Total volume by data type, dialect mix with percentage targets, quality thresholds (IAA by task), annotator qualification requirements, PDPL compliance documentation, delivery format and schema, timeline, and a mandatory pilot before volume commitment.

What is the difference between SFT and RLHF data for Arabic LLMs?

SFT (supervised fine-tuning) data consists of instruction-response pairs that teach the model target behaviour. RLHF (reinforcement learning from human feedback) data consists of preference comparisons — ranked model outputs — used to train a reward model. Both require native-speaker raters by dialect; RLHF additionally requires raters who can judge helpfulness and cultural appropriateness.

How much does Arabic LLM training data cost?

Pricing varies by task complexity, dialect, and volume. General-knowledge MSA SFT pairs are at the lower end; Khaleeji-native domain-expert RLHF comparisons are at the higher end. See our pricing page or request a project quote for your specific dialect mix and volume.

What PDPL requirements apply to Arabic LLM data?

If your training data contains personal information about Saudi residents, Saudi PDPL requires a data processing agreement, a lawful transfer mechanism for cross-border movement, and documented deletion after use. Confirm DPA, transfer mechanism, and deletion procedures with any vendor before signing.

How do you evaluate an Arabic LLM data vendor?

Ask for a dialect pool breakdown with annotator counts, IAA scores from comparable recent projects, a sample annotation guideline document, a PDPL-compliant DPA template, and a pilot of 50–100 pairs before volume commitment. Reject any vendor who cannot produce per-dialect IAA reports.

Free Sample · 24-48 hours

Request Arabic LLM Training Data

Tell us your dialect mix, volume, and task types. We'll respond with a project brief and a free 50-pair pilot within 5 business days.

This form is for companies with annotation projects. Looking for annotation work? Apply on our careers page. Job enquiries sent here don't get a reply.

No commitment. NDA available on request. We respond within 24 hours, often the same day for Gulf-region inquiries.

Neel Bennett

AI Annotation Specialist at AI Taggers

Neel has over 8 years of experience in AI training data and machine learning operations. He specializes in helping enterprises build high-quality datasets for computer vision and NLP applications across healthcare, automotive, and retail industries.

Connect on LinkedIn