AuthorityCompliance

The EU AI Act and Your Training Data: What Actually Changes

The EU AI Act's high-risk provisions apply from August 2026. For AI teams building systems in healthcare, finance, employment, or critical infrastructure, the training data requirements in Article 10 are not a future compliance checkbox — they are active obligations now. Here is what they require, who they apply to, and what the annotation supply chain needs to change.

1 October 202613 min read

Direct answer

The EU AI Act (Regulation 2024/1689, fully applicable for high-risk AI from August 2026) requires that training data for high-risk AI systems be relevant, representative, and as free of errors and complete as possible — and that the data governance practices used to collect, process, and annotate that data be documented. For annotation teams, this means three concrete changes: provenance records for every dataset used in high-risk training; bias examination documentation, particularly for demographic coverage; and quality assurance audit trails sufficient to demonstrate that annotation errors were actively managed rather than passively tolerated.

The Regulatory Context: What the EU AI Act Is and Who It Reaches

The EU AI Act (Regulation 2024/1689) is the first comprehensive binding legal framework for artificial intelligence in any major jurisdiction. It entered into force in August 2024, with a phased application timeline. Prohibited AI practices applied from February 2025. General-purpose AI model (GPAI) obligations applied from August 2025. The most operationally significant tier — high-risk AI system requirements including Article 10 training data obligations — applies from August 2026.

The Act has extra-territorial scope comparable to GDPR. It applies to any provider placing an AI system on the EU market, regardless of where the provider is established. An Australian company whose credit-scoring model processes EU residents' applications, or whose recruitment AI screens candidates for EU-based roles, falls within scope. The European AI Office estimates that the Act affects approximately 50,000 deployers and 5,000 providers of AI systems across the EU and in third countries serving the EU market (European Commission, 2024 impact assessment).

The compliance obligation sits primarily with the AI system provider — the entity that develops and places the system on the market. But providers are responsible for their entire data supply chain. Annotation vendors supplying training data for high-risk AI systems are the upstream link in that chain.

What Article 10 Actually Requires for Training Data

Article 10 is the central training data provision for high-risk AI. Its requirements can be grouped into four areas:

1. Data quality requirements

Training, validation, and test data must be "relevant, sufficiently representative, and, to the best extent possible, free of errors and complete with regard to the intended purpose." The Act specifically requires that data practices take into account the characteristics of the specific geographic, contextual, behavioural, or functional setting in which the AI system is intended to be used.

In practice, this means that a recruitment AI deployed across EU member states cannot be trained exclusively on historical data from one country — the training data must reflect the demographic and cultural range of the deployment context. For annotation teams, it means that coverage gaps in the training data (limited data from specific age groups, nationalities, or socioeconomic contexts) must be identified and documented rather than ignored.

2. Data governance documentation

Article 10(2) requires data governance practices to address: the design choices about data collection, data origins, the data processing steps and their rationale (pre-processing, annotation, labelling, cleaning, enrichment, and aggregation), any known limitations and how these are addressed, the purposes of the processing, and any examination for possible biases that could affect health and safety or lead to prohibited discrimination.

This is not a vague documentation requirement. It is a specific list of what must be recorded. "Data origins" means data source, collection method, and jurisdiction. "Annotation" means the methodology, the inter-annotator agreement metrics, and the QA process. "Known limitations" means actively assessing and disclosing dataset coverage gaps, not asserting completeness.

3. Bias examination

Article 10(2)(f) explicitly requires examination of the training data for possible biases "that could affect the health and safety of natural persons or lead to discrimination prohibited by Union law, in particular where data outputs influence inputs for future operations." Prohibited discrimination under EU law covers race, ethnic origin, religion, disability, age, sexual orientation, sex, and other characteristics in the EU Charter of Fundamental Rights.

This requirement creates a concrete obligation: high-risk AI developers must be able to demonstrate that they examined their training data for demographic and representational bias, and either addressed it or documented why they could not. A recruitment model trained on historical hiring data without a documented bias examination does not meet this requirement.

For annotation teams, this translates into a requirement to maintain demographic coverage records for annotated datasets, document known coverage gaps, and where possible provide bias metrics (annotator demographic distribution, label distribution by demographic subgroup where ethically permissible) to the AI system provider for their compliance documentation.

4. Special category data

Article 10(5) permits the processing of special category personal data (biometric, health, racial or ethnic origin, etc.) for the purpose of bias detection and correction, subject to appropriate safeguards. This is a limited permission that does not override GDPR — it permits processing that GDPR would otherwise restrict, but only for the specific purpose of bias examination within the AI system development lifecycle.

Need EU AI Act-ready annotation documentation?

Our data QA and validation service provides the provenance records, inter-annotator agreement metrics, and bias examination documentation that Article 10 compliance requires.

Talk to our compliance-aware annotation team

Which AI Systems Fall Into the High-Risk Category

Annex III of the EU AI Act lists the high-risk AI categories. The breadth is wider than most teams anticipate:

Medical device AI is not in Annex III — it is governed by the Medical Device Regulation (MDR) and In Vitro Diagnostic Regulation (IVDR), which have analogous and in some cases more stringent training data requirements. The EU AI Act creates an additional layer of obligations for AI components within medical devices that meet the AI definition.

Case Study: Annotation Documentation for an EU Credit-Scoring Model

A European fintech was building an automated credit risk model for use across Germany, France, and Poland. The model used a combination of structured financial data and NLP analysis of customer-submitted documents — income verification letters, bank statements, and employer confirmations. The NLP component required annotation of document entity extraction and risk-relevant text classification.

Their initial annotation process was effective from a quality standpoint: inter-annotator agreement on the primary extraction task reached 0.89 kappa, and model performance on held-out validation sets was strong. But their documentation was sparse. When their legal team conducted an Article 10 readiness assessment in early 2026, the gaps were significant: no record of annotator demographic distribution, no documentation of how annotators were trained to handle the three national document formats, no coverage analysis across the three deployment markets, and no explicit bias examination for whether the annotation scheme itself created systematic differences in how documents from different demographic groups were labelled.

The remediation required a retrospective documentation effort plus a prospective annotation redesign. The annotation methodology was restructured to include: annotator origin and linguistic background records (German, French, and Polish native speakers distributed across the annotation pool), document-type coverage sampling to ensure each national document format was represented in training data at the frequency it appears in the deployment market, and explicit bias review sessions where senior annotators reviewed a stratified sample of annotations for systematic differential treatment by document origin.

The bias review found one genuine issue: annotator confidence in labelling Polish document dates was lower (reflected in higher correction rates) because Polish date formats were under-represented in the annotation guidelines examples. This had produced a subtle but measurable difference in extraction accuracy for Polish-origin documents — 91.3% vs 96.1% for German-origin documents on the same entity type. The guidelines were updated with explicit Polish format examples, and the affected records were relabelled. Both the issue and the remediation were documented for the Article 10 compliance file.

What Changes for Annotation Vendors and Buyers

The Article 10 obligations sit with the AI system provider, but the documentation requirements flow upstream. In practice, AI developers building high-risk systems need their annotation vendors to provide records that make compliance documentation possible. This creates several concrete changes in how annotation engagements are structured:

Provenance records as a deliverable

The annotation output is no longer just labelled data. It includes methodology documentation: where the source data originated, how annotators were selected and trained, what the inter-annotator agreement metrics were, how disagreements were resolved, and what the known coverage limitations are. This is documentation that data QA and validation processes must generate systematically, not retrospectively.

Bias examination as a project phase

Bias examination can no longer be an afterthought. For high-risk AI, annotation projects need a planned bias review phase: stratified sampling by relevant demographic dimensions, comparison of annotation quality metrics across demographic subgroups, and documented findings that go into the provider's Article 10 compliance file. Our annotation QA and relabelling capability includes structured bias review sessions for regulated AI projects.

Contractual alignment

EU AI Act compliance creates contractual obligations between AI system providers and their data supply chain. Annotation vendors working on high-risk AI projects should expect to see data processing agreements that include specific requirements around documentation retention, audit rights, and notification of methodology changes. This mirrors the evolution of GDPR data processor agreements — and in many cases will be handled within the same DPA framework, since annotation of personal data almost always involves GDPR obligations as well.

The GPAI Tier: What Changes for LLM Training Data

General-purpose AI models (GPAI models) — including large language models — are covered under Title V of the Act, which applied from August 2025. The GPAI obligations differ from the high-risk AI Article 10 requirements in important ways.

For standard GPAI models (below the systemic-risk threshold of approximately 10^25 training FLOPs), providers must maintain technical documentation that includes: the training data used (including type, format, provenance), the data governance practices applied, measures to comply with EU copyright law regarding training data (Article 53(1)(c) — this is the provision that requires GPAI providers to have a policy on third-party rights and copyright compliance, and to make a sufficiently detailed summary of training data publicly available).

For GPAI models with systemic risk (the largest frontier models), additional obligations apply: adversarial testing, incident reporting to the European AI Office, and cybersecurity measures. The training data documentation requirements are substantively the same — the extra tier adds testing and monitoring obligations, not fundamentally different data governance requirements.

Practical Implications for Australian AI Teams

Australia does not have an equivalent of the EU AI Act — the Australian AI Ethics Framework is voluntary, and the government's 2024 regulatory consultation resulted in a proposal for mandatory guardrails for high-risk AI but has not yet been legislated as of October 2026. However, Australian AI teams face EU AI Act obligations in two common scenarios:

For Australian AI developers in healthcare, financial services, HR technology, and critical infrastructure, the EU market is sufficiently significant that building to EU AI Act standards from the start is prudent — rather than maintaining separate annotation documentation processes for EU and non-EU use cases.

The good news: the documentation requirements of Article 10 align closely with best-practice annotation methodology that good annotation vendors should already be providing. If your annotation workflow produces inter-annotator agreement metrics, provenance records, and systematic QA audit trails, the marginal effort to produce Article 10-compliant documentation is modest. If it does not, both your annotation quality and your regulatory standing need attention — and the former is likely costing your model performance regardless of the regulatory question.

Building a Compliance-Ready Annotation Pipeline

If you are preparing for EU AI Act compliance on the training data side, the changes are structural rather than cosmetic. The following are the highest-priority elements of a compliance-ready annotation pipeline:

These are not separate compliance deliverables — they are what a well-run annotation process should produce as standard outputs. The EU AI Act did not invent best-practice annotation; it mandated documentation of practices that responsible annotation has always required. You can learn more about how our data QA and validation process produces these records, or read our guide on writing annotation guidelines that produce consistent, documentable output.

Frequently Asked Questions

What does the EU AI Act require for training data?

For high-risk AI systems (Article 10), the EU AI Act requires training data that is relevant, representative, free of errors, and complete to the extent possible, plus documentation of data governance practices covering collection methodology, data origin, processing steps, known limitations, and bias examination. These requirements apply to the provider of the high-risk AI system.

What counts as a high-risk AI system under the EU AI Act?

Annex III lists the high-risk categories: biometric identification, critical infrastructure AI, education AI, employment AI (recruitment, monitoring, performance), essential services AI (credit scoring, insurance), law enforcement AI, migration and border AI, and AI in justice administration. Medical device AI is covered by MDR/IVDR but the Act creates additional obligations for AI components.

Does the EU AI Act apply to AI systems outside the EU?

Yes. The Act applies to providers placing AI systems on the EU market regardless of where the provider is established, and to deployers operating within the EU. An Australian company whose AI is used by EU citizens or deployed in the EU market falls within scope.

What is the EU AI Act compliance deadline for training data?

High-risk AI system obligations, including Article 10 training data requirements, apply from August 2026. Systems deployed before the Act's application date have until August 2027. GPAI model obligations (including training data documentation) applied from August 2025.

What documentation does an annotation vendor need for EU AI Act compliance?

Annotation vendors supplying training data for high-risk AI should provide: data provenance documentation, annotation methodology and IAA metrics, bias examination records, known limitations and coverage gaps, and QA audit trails. AI system providers building Article 11 Technical Documentation will request this from their annotation suppliers.

Does the EU AI Act affect LLM training data?

Yes, under the GPAI provisions (Title V, applied August 2025). GPAI providers must document training data type, format, and provenance, and publish a summary of training data. For systemic-risk GPAI models (above ~10^25 training FLOPs), additional testing and incident reporting obligations apply.

Free Sample · 24-48 hours

Build a compliance-ready annotation pipeline

We provide provenance records, IAA metrics, bias examination documentation, and QA audit trails that satisfy EU AI Act Article 10 requirements for high-risk AI training data.

No commitment. NDA available on request. We respond within 24 hours, often the same day for Gulf-region inquiries.

Neel Bennett

AI Annotation Specialist at AI Taggers

Neel has over 8 years of experience in AI training data and machine learning operations. He specializes in helping enterprises build high-quality datasets for computer vision and NLP applications across healthcare, automotive, and retail industries.

Connect on LinkedIn