A data annotation RFP (Request for Proposal) is a procurement document that specifies annotation task requirements, quality expectations, compliance requirements, and evaluation criteria to invite structured vendor responses. An effective annotation RFP has seven required sections: project overview, task specification with examples, dataset description, quality and QA requirements, compliance and security requirements, commercial terms, and evaluation scoring criteria. The most effective RFPs also include a paid pilot requirement — pilots of 500–1,000 records are the only reliable way to verify vendor quality claims before awarding a full contract.
Why Most Annotation RFPs Fail — And What to Do Instead
A 2025 Gartner survey of enterprise AI procurement teams found that 58% of data annotation vendor selections resulted in at least one project requiring significant rework or vendor replacement within the first six months. In most cases, the vendor was selected through an RFP process — the failure was not in the vendor evaluation but in what the RFP asked and how it weighted responses.
The three most common RFP failure modes in annotation procurement are: under-specifying the task (vendors cannot price accurately without a real annotation schema and examples), evaluating primarily on per-label price (the cheapest quoted rate rarely produces the lowest total delivery cost), and omitting compliance requirements from the initial RFP (discovering incompatibility with your data handling requirements after vendor selection typically delays projects by 4–8 weeks).
This guide gives you the section structure, required content, scoring framework, and pilot design that produce reliable vendor selection outcomes. For a detailed look at how annotation pricing structures affect total delivery cost — which directly informs how to evaluate RFP responses — see our guide to annotation pricing models.
Section 1: Project Overview and Business Context
The project overview should give vendors enough business context to understand the downstream use of the annotation output. Vendors who understand the end application — a fraud detection model, a medical diagnosis AI, an autonomous vehicle perception system — make better annotation decisions at the margin and write more realistic proposals.
Required content for the project overview section:
- Organisation and project name — identifies the buyer and creates a named context for vendor questions
- Business purpose — what the annotated dataset will train or evaluate, without requiring proprietary detail
- Project scope summary — estimated total label volume, task types, and dataset source (internal data, third-party purchase, synthetic)
- Key constraints — hard deadlines, budget envelope if disclosed, jurisdictional data residency requirements
- Decision timeline — RFP response deadline, shortlist notification date, pilot start date, and target contract award date
A common mistake in this section is describing the annotation task as a generic type ("image annotation") rather than the specific application ("bounding box annotation for surgical instrument detection in laparoscopic video"). The specific application determines which vendor capabilities are relevant and allows vendors to provide realistic pricing based on comparable project experience.
Section 2: Task Specification — The Make-or-Break Section
The task specification section is where most annotation RFPs under-deliver. Without a real annotation schema, example items, and edge case guidance, vendor pricing is speculative and vendor capability claims are unverifiable. The best task specifications include draft annotation guidelines — even a two-page version — that vendors can use to assess task complexity and price accurately.
Minimum required content for task specification:
- Annotation type — bounding box, semantic segmentation, NER, sentiment classification, transcription, etc.
- Label schema — the complete list of classes, attributes, and annotation rules, with at least 3 example items per class
- Input data format — file types, resolutions, languages, audio sample rates, or document structures as applicable
- Output format requirement — COCO JSON, YOLO, PASCAL VOC, CSV, custom schema
- Known edge cases — the 3–5 edge case types you already know exist in the dataset
- Annotator expertise required — general-domain, domain-expert (clinical, legal, financial), or bilingual/native speaker
For complex or specialised tasks, attach a sample dataset of 20–50 items that vendors can use for pilot pricing and capability assessment. A sample dataset reduces pricing variance across RFP responses, which makes evaluation more meaningful. For custom or non-standard annotation workflows, our custom annotation service page describes how to scope non-standard tasks for RFP purposes.
Section 3: Quality Requirements
Quality requirements are the most commonly vague section of annotation RFPs — and the most important. "High quality" and "99% accuracy" are not specifications; they are aspirations. A well-written quality requirements section defines quality precisely enough that both buyer and vendor can verify compliance objectively.
Required quality specification elements:
- Accuracy threshold — minimum acceptable accuracy rate (e.g., ≥97% on gold-standard test items) or minimum IAA (e.g., Cohen's κ ≥ 0.80 for classification tasks)
- QA methodology — how quality will be measured: expert review sample rate, gold-standard injection rate, inter-annotator agreement measurement
- Rework policy — define who funds rework of labels falling below the quality threshold
- Escalation process — how edge cases or ambiguous items should be handled and by whom
- Reporting requirements — what quality metrics are reported, at what frequency, and in what format
Ask vendors to describe their QA process in detail in their response — specifically: what percentage of labels are reviewed, by whom, using what method. A vendor who cannot describe their QA process in operational detail is unlikely to have one that will meet your quality requirements. See our guide to the annotation QA process for the full framework of quality measurement methods to reference in your RFP.
Want Help Scoping Your Annotation RFP?
We respond to annotation RFPs with detailed project proposals, transparent pricing across all three billing models, and reference case studies for your specific task type.
See Custom Annotation ServicesSection 4: Compliance and Security Requirements
Compliance requirements should be specified upfront in the RFP — not surfaced after vendor selection. Including a compliance checklist in the RFP allows vendors to self-screen for eligibility and prevents a common procurement failure: selecting a vendor at a competitive rate and discovering post-selection that they cannot meet your organisation's data handling requirements.
Key compliance elements to specify for Australian enterprise projects:
- Data sovereignty — whether data must remain in Australia or a specified jurisdiction (relevant for government and healthcare buyers under the Australian Privacy Act 1988)
- Security certifications — ISO 27001, SOC 2 Type II, or equivalent information security certifications required
- Industry-specific compliance — HIPAA for US-market healthcare data, FDA 21 CFR Part 11 for clinical trial data, PDPL for Saudi Arabian market data
- NDA and IP assignment — whether annotated datasets require IP assignment to buyer, and whether vendor sub-contracting is permitted
- Annotator vetting — background check requirements, confidentiality agreement requirements for annotators touching sensitive data
- Breach notification — contractual obligation for data breach notification within defined timeframes
For a detailed guide to compliance requirements across GDPR, PDPL, and HIPAA as they apply to annotation outsourcing, see our post on GDPR vs PDPL vs HIPAA compliance for annotation buyers.
Section 5: Commercial Terms and Pricing Model Preference
Specify which pricing model you prefer or invite vendors to propose alternatives. Requiring vendors to price in per-label, hourly, and managed-service terms allows side-by-side comparison — but also reveals which model each vendor's business is optimised for. A vendor who quotes per-label rates 40% above market for a straightforward classification task but is competitive on managed service is telling you something about their operational model.
Required commercial information to request:
- Per-label, hourly, or managed-service quote for the specified scope (request all three if undecided on model)
- Rework policy — whether rework is included in quoted rate or billed additionally
- Minimum project size and ramp-up timeline
- Volume discount tiers if applicable
- Payment terms preference
- Any additional fees — tooling, setup, project management, reporting
Require vendors to provide a total cost of delivery estimate, not just a per-unit rate. The total cost of delivery estimate should include their projected rework rate and any tooling or setup fees — this is the number you compare across vendors, not the headline per-label rate. For the full breakdown of how to compare annotation costs including internal QA overhead, see our annotation pricing models guide.
Vendor Scoring Framework: The 5-Category Rubric
Use a weighted scoring rubric to evaluate vendor responses consistently across your shortlist. The following weightings work well for most enterprise annotation procurement decisions:
- Technical capability (25%) — demonstrated experience with your specific task type, annotator expertise available, tooling capabilities
- Quality processes (25%) — QA workflow detail, IAA measurement methodology, rework track record from references
- Security and compliance (20%) — certifications held, data handling practices, sub-contracting policies
- Pricing and commercial terms (20%) — total cost of delivery estimate (including rework and management overhead), payment terms, contract flexibility
- References and case studies (10%) — verifiable experience with similar project type, scale, and domain
Score each vendor on a 1–5 scale for each criterion, multiply by the weighting, and sum for a total score. This approach prevents a low price from dominating the decision when quality processes are materially weaker — which is the most common source of annotation vendor selection regret.
Note that the scoring rubric weights are starting points, not fixed values. For high-compliance projects (healthcare, government, legal), increase security/compliance to 30% and reduce pricing accordingly. For commodity annotation tasks with strong vendor supply, price weighting can increase to 30% without material risk.
Project Case Study: RFP That Saved 14 Weeks and AUD $340K
An Australian health technology company was procuring annotation services for a 600,000-image medical imaging dataset. Their initial RFP was two pages long: it described the task as "medical image annotation" with an accuracy requirement of "95% or better" and asked for per-image pricing.
Six vendors responded. The lowest-price vendor was selected at AUD $1.20 per image. At the four-week mark, the internal QA team was rejecting 28% of delivered labels — primarily because the RFP had not specified the annotation schema, and the vendor had interpreted the task differently from the buyer's intent. The project was paused at 42,000 images — AUD $50,400 spent — with no usable dataset.
The buyer re-ran the procurement using a structured RFP with the full task specification (detailed schema, 50 example images, five documented edge cases), a compliance checklist, and a mandatory pilot requirement. The pilot was sent to three shortlisted vendors — 500 images each, paid. Two vendors achieved 96–98% agreement with the buyer's gold standard; one achieved 71% and was eliminated.
Final outcomes compared to the first procurement attempt:
- Selected vendor: AUD $1.75/image (46% higher than original vendor rate)
- Delivered dataset rejection rate: 2.8% (vs 28% in the first attempt)
- Project completed: 11 weeks (vs an estimated 28-week trajectory under the original vendor)
- Total cost saving (including rework avoided and engineering time): approximately AUD $340,000
The structured RFP and pilot requirement added approximately three weeks to the procurement timeline. Those three weeks recovered 14 weeks of project timeline and AUD $340K in total delivery cost. For help scoping the task specification section of your RFP — the section most likely to determine vendor selection quality — see our annotation project scoping checklist.
The Pilot Project: Your Best Vendor Evaluation Tool
For any project over AUD $50,000 in annotation spend, a paid pilot is the most valuable investment in the procurement process. No RFP response can demonstrate what 500–1,000 annotated records from your actual dataset can.
A well-designed annotation pilot reveals:
- Actual quality on your data — not quality claims, not reference project quality, but actual inter-annotator agreement on your specific annotation schema
- Throughput on your task — the benchmark for project timeline estimation and cost modelling
- Guideline interpretation — how the vendor handles ambiguous cases reveals their annotator training, escalation processes, and communication habits
- Communication responsiveness — a vendor who is slow to respond during a paid pilot, when they are trying to win the contract, will be slower during production
- Tooling and file delivery — the format, completeness, and metadata of pilot deliverables tells you what production deliverables will look like
Design pilots to be representative, not cherry-picked. Include known-difficult edge cases from your dataset in the pilot sample — the goal is to stress-test vendor capability, not to make the pilot easy enough that every vendor passes.
Pay vendors for pilots at their standard rate. Unpaid pilots bias responses toward vendors with spare capacity rather than the most capable vendors. A AUD $1,000–$3,000 pilot investment for a project worth AUD $200,000+ is among the highest ROI procurement activities available. Our custom annotation service includes structured pilot scoping as part of every initial project engagement.
Questions to Ask Shortlisted Vendors in RFP Clarification
The written RFP response tells you what vendors want you to hear. Clarification questions reveal how they actually operate. Ask every shortlisted vendor:
- "Describe the exact QA process you use for a project like ours — what percentage of labels are reviewed, by whom, and what happens when a batch fails QA?"
- "What was the rework rate on your three most recent similar projects?"
- "How do your annotators handle edge cases not covered by the guidelines?"
- "What is your annotator turnover rate, and how do you handle knowledge transfer mid-project?"
- "Can you provide the contact details of a project manager from a reference project we can speak with directly?"
- "What is your process when a client's requirements change mid-project?"
Vendors who give vague or scripted answers to these operational questions, or who are reluctant to provide direct client references, are surfacing risk that the written RFP response concealed. The clarification call is your opportunity to distinguish vendors who have mature operational processes from vendors who have good proposal writing.
FAQ
What should a data annotation RFP include?
A data annotation RFP should include: project overview and business context, task specification with schema and examples, dataset description, quality requirements with measurable thresholds, compliance and security requirements, commercial terms including pricing model preference, and evaluation scoring criteria with weightings. The most commonly missing section is compliance — discovering incompatibility after vendor selection typically delays projects by 4–8 weeks.
How do you score annotation vendor RFP responses?
Use a weighted rubric: technical capability (25%), quality processes (25%), security and compliance (20%), pricing and commercial terms (20%), references and case studies (10%). Score each vendor 1–5 per criterion, multiply by weighting, sum for a total. Avoid over-weighting price — the cheapest per-label rate rarely produces the lowest total cost of delivery.
How long should a data annotation RFP process take?
A well-run process takes 6–10 weeks: 1–2 weeks to write and distribute the RFP, 2–3 weeks for vendor responses, 1–2 weeks for shortlisting and clarification, 1–2 weeks for pilot projects, and 1 week for contract negotiation. Cutting corners on the pilot phase is the most common cause of vendor selection failures.
Should you require a pilot project in an annotation RFP?
Yes — for any project over AUD $50,000 or involving complex tasks. A paid pilot of 500–1,000 records reveals quality, throughput, communication responsiveness, and guideline interpretation that no written RFP response can demonstrate. Shortlist 2–3 vendors for parallel pilots; the winner is the vendor whose pilot quality and working style best match your project needs.
What are the most common data annotation RFP mistakes?
The five most common: under-specifying the task (no schema or examples); omitting compliance requirements; evaluating primarily on per-label price; skipping pilot projects; and writing a generic RFP rather than one tailored to your specific task type. Medical, legal, and multilingual projects require task-specific evaluation criteria that generic templates do not cover.
How many vendors should you include in an annotation RFP?
Send the RFP to 4–6 shortlisted vendors. Shortlist from 8–12 vendors based on task-type fit, geography and time zone alignment, and compliance certifications. Running pilots with more than 3 vendors adds evaluation cost without proportional decision-making benefit.
Need Help Writing Your Annotation RFP?
We respond to annotation RFPs with detailed proposals including task-specific pricing, QA process documentation, compliance certifications, and reference case studies.
Neel Bennett
AI Annotation Specialist at AI Taggers
Neel has over 8 years of experience in AI training data and machine learning operations. He specializes in helping enterprises build high-quality datasets for computer vision and NLP applications across healthcare, automotive, and retail industries.
Connect on LinkedIn