Direct answer
Data-centric AI is an approach to machine learning that prioritises systematic improvement of training data quality over model architecture iteration. The core evidence: Andrew Ng's 2021 experiments showed that consistent data quality improvement outperformed architecture switching on a speech recognition task, producing twice the accuracy gain at lower cost. For most production AI tasks with datasets under 100k examples, this pattern holds — the model is not the bottleneck; the data is.
The Shift That Changed How Production AI Teams Operate
For the first decade of the deep learning era, the dominant frame was model-centric: hold the dataset fixed and iterate on the model. The ImageNet competition rewarded architecture innovation. Benchmarks were fixed datasets. The field's metrics of progress were model accuracy on standard held-out sets, not data quality.
This made sense for academic research where the benchmark is the goal. It does not make sense for production AI, where the goal is a deployed system that performs correctly on real-world inputs. In production, the dataset is not fixed — it is collected, annotated, and continuously updated. And in practice, the annotation process is rarely managed with the same engineering rigour as the model training process.
The result: enterprise AI teams spending months iterating on model architecture while their training data contains systematic label errors, coverage gaps, and distribution mismatches that no model can compensate for. Data-centric AI names this failure mode and provides a corrective framework.
The Evidence: What Ng's Experiments Actually Showed
In 2021, Andrew Ng ran a structured experiment to compare model-centric and data-centric improvement strategies on a speech recognition task. The setup: a fixed baseline model, a fixed test set, and two improvement paths — one that iterated on model architecture and hyperparameters while holding data constant, one that systematically improved data quality while holding the model constant.
The data-centric path achieved approximately 2× the accuracy improvement of the model-centric path. The specific intervention was consistent labelling: ensuring that similar utterances were labelled consistently across the training set, that noise and artefacts were handled uniformly, and that annotation guidelines produced the same decisions for similar inputs. No new data was collected — existing data was relabelled more carefully.
This experiment was influential not because the result was surprising in isolation — practitioners had long observed that data quality mattered — but because it was structured, comparable, and public. It gave the data quality observation the empirical backing needed to change how teams allocate engineering effort.
Data improvement vs model improvement: typical outcomes
Source: Internal project benchmarks across 30+ enterprise annotation projects, 2024–2026. Ranges vary by task, dataset size, and baseline data quality.
What Data-Centric AI Means in Practice
Data-centric AI is not a specific algorithm or framework — it is a posture: a set of priorities and practices that treat data improvement as the primary lever for model improvement. In practice, it involves three interlocking activities:
1. Systematic data quality measurement
You cannot improve what you cannot measure. Data-centric AI programmes begin with a structured quality audit: measuring inter-annotator agreement (IAA) to quantify labelling consistency, constructing gold-standard sets to measure annotator accuracy, and profiling error rates by label class and data subtype to identify where quality is degrading model performance.
The Cohen's kappa coefficient is the standard IAA metric for categorical annotation tasks. A kappa below 0.7 indicates a labelling consistency problem at the guidelines level. A kappa above 0.85 indicates the annotation task is well-defined and consistently executed. These thresholds give teams an objective signal for when to invest in annotation quality improvement versus when to move to model iteration.
2. Targeted data improvement, not blanket re-annotation
A common misreading of data-centric AI is that it means re-annotating everything. It does not. The most efficient data-centric programmes identify the specific slices of the training data that are most affecting model performance — a particular label class, a specific data source, a distribution gap in the training set — and direct improvement effort there.
This is where active learning and human-in-the-loop workflows become powerful: the model itself surfaces the examples it is most uncertain about, directing annotator attention to the data that will have the highest marginal impact on accuracy. When the active learning loop is correctly implemented, each annotation cycle produces compounding quality improvements rather than linear ones.
AI Taggers' custom annotation workflows are designed to integrate this kind of targeted improvement — starting from a data audit and prioritising relabelling effort based on measured impact, not uniform coverage.
3. Annotation infrastructure that supports iteration
Data-centric AI is iterative. Each improvement cycle produces a new training set, a new model, and new performance measurements that feed the next audit. This requires annotation infrastructure that supports version control of datasets, reproducible measurement of IAA, and systematic tracking of which label schema versions produced which model performance.
Many organisations do not have this infrastructure. They annotate once, train once, and discover problems in production. Data-centric AI requires a different operational model: annotation as an ongoing engineering activity, not a one-time procurement.
Ready to take a data-centric approach to your AI project?
AI Taggers designs custom annotation workflows built around iterative data quality improvement — from initial audit to production-grade pipeline.
See our custom annotation servicesA Real Data-Centric Programme: NLP Classifier from 71% to 94% F1
An Australian insurtech company engaged AI Taggers to improve a claims-routing NLP classifier that had plateaued at 71% F1 on a 12-class routing taxonomy. Their ML team had run eight model iterations over six months — BERT, RoBERTa, domain-adapted variants — without breaking 75%. The engineering lead suspected a data problem but had no measurement infrastructure to confirm it.
The data-centric programme proceeded in three phases:
Audit phase (2 weeks)
Gold-set audit on 800 training records revealed a 14.7% label error rate overall, concentrated in three routing categories where the taxonomy had changed mid-project and old labels had not been updated. IAA measurement on a 200-record sample returned kappa of 0.61 — well below acceptable threshold. Two annotation guidelines documents were in use simultaneously, creating systematic annotation divergence.
Remediation phase (4 weeks)
Consolidated guidelines with 25 concrete decision-boundary examples per high-confusion class pair. Relabelled 6,200 records in the three high-error categories. Removed 340 records where claims text had been partially redacted, making label assignment ambiguous — these had been contaminating training since project inception. Re-measured IAA: kappa 0.86.
Expansion phase (3 weeks)
Active learning run on the updated model surfaced 1,400 low-confidence examples concentrated in two routing classes with thin training coverage. Annotated these at double density (three annotators, adjudication on disagreements). Final retrain on same RoBERTa architecture used for the original baseline.
Results: same model, systematically improved data
The model that had plateaued at 71% for six months of architecture iteration reached 94% in nine weeks of data improvement. The architecture did not change. The data did. This is data-centric AI operating as intended.
Where Data-Centric AI Has the Most Impact
Data-centric approaches produce the largest gains in specific contexts. Understanding where the lever is largest helps teams prioritise investment:
Small to medium datasets (under 100k examples)
At small dataset scale, label errors represent a higher fraction of the total signal and the model has less opportunity to average out noise. Each mislabelled example has disproportionate influence. Data-centric improvement here is high-leverage.
Complex or ambiguous label schemas
When the annotation task requires nuanced judgement — subtle sentiment, overlapping entity types, multi-label classification with correlated categories — annotation consistency is intrinsically harder to achieve and intrinsically more impactful when achieved. Tasks like claims routing, medical coding, and moderation have this character.
Production systems that have drifted from their training distribution
Models trained 12–24 months ago on historical data increasingly diverge from the current production distribution. Expanding training data coverage to match current production inputs — a data-centric intervention — can recover performance without retraining from scratch.
Foundation model fine-tuning
Fine-tuning a large pre-trained model is almost entirely a data-centric problem. The base model has broad capability; the question is what examples to fine-tune it on. Quality of the fine-tuning dataset — typically hundreds to thousands of examples — dominates quality of the fine-tuned model.
High-stakes domains: medical, financial, legal
When annotation requires domain expertise — radiologists labelling pathology, clinicians annotating clinical text, compliance specialists reviewing financial documents — annotator quality variation is large and the cost of systematic errors is high. Data-centric principles, applied to expert annotation pipelines, produce large quality improvements at these quality tiers.
Data-Centric AI and Custom Annotation Workflows
The operational implication of data-centric AI is that annotation cannot be a commodity procurement: buy labels, run training, deploy model. It requires custom annotation workflows that are designed around the specific quality requirements of each project — with measurement built in from the start, not audited at the end.
A data-centric annotation workflow typically includes:
- Pilot annotation with IAA measurement before full-scale production, to validate guidelines and identify ambiguities at low cost
- Gold sets embedded in annotation batches throughout the project, providing continuous quality monitoring rather than end-of-project audits
- Slice-based quality reporting — performance broken down by label class, data source, and input subtype — so quality problems are visible before they propagate into training
- Active-learning integration to direct annotation effort toward the inputs with the highest marginal impact on model performance
- Version-controlled label schemas so that guideline changes are tracked, their impact on training data is known, and historical labels can be reviewed against updated standards
This infrastructure is not optional overhead — it is what makes the data-centric improvement loop possible. Without it, each annotation cycle is disconnected from the previous one, and the compounding quality improvements that data-centric AI promises are not achievable.
Applying Data-Centric Principles to Build vs Buy Decisions
Data-centric AI has direct implications for the build vs buy annotation decision. In-house annotation teams are often constructed around the model-centric assumption: label a dataset once, use it, move on. The data-centric model requires ongoing annotation capability — an audit-remediate-retrain cycle that does not have a natural end point.
For most organisations, sustaining the data-centric loop in-house is feasible for a single high-priority project but difficult across a portfolio of AI initiatives. External annotation partners who have invested in measurement infrastructure, expert annotator pools, and IAA-based QA protocols can run the improvement cycle at lower cost than in-house teams while producing higher-quality data.
The evaluation criterion shifts accordingly: the right annotation partner for a data-centric programme is not necessarily the cheapest per label, but the one whose pipeline produces the highest measurable data quality and whose reporting infrastructure makes the improvement cycle visible. For a transparent breakdown of what production-quality annotation costs by task type, see our 2026 annotation pricing guide.
Data-Centric AI in 2026: Where the Field Has Settled
Five years after Ng's initial 2021 campaign, data-centric AI has moved from a corrective framing to a mainstream practice in enterprise ML. The major shifts:
Foundation model fine-tuning has made data quality even more central. When the base model is pre-trained at massive scale, the differentiation between products built on it is almost entirely in fine-tuning data quality. RLHF and DPO training for alignment — preference data annotation — is a canonical data-centric activity where annotation quality directly determines whether the fine-tuned model is helpful or harmful.
Tooling for data quality measurement has matured. CleanLab and Confident Learning from MIT formalizsed the algorithmic detection of label errors at scale. Label Studio, Argilla, and enterprise platforms now include IAA measurement as standard features. The measurement infrastructure that data-centric AI requires is no longer bespoke engineering.
The "synthetic data will replace annotation" thesis has been tested and found partial. Synthetic data is effective for specific tasks — object detection augmentation, test set expansion, edge-case generation — but systematic tests have shown that synthetic data for NLP fine-tuning consistently underperforms real human-annotated data when the task requires genuine language understanding. The synthetic vs annotated data question is now well-characterised: complementary rather than substitutable for most production tasks.
Frequently Asked Questions
What is data-centric AI in simple terms?▼
Is data-centric AI the same as data quality?▼
How do you know when to focus on data vs model?▼
Can data-centric AI work for LLM fine-tuning?▼
What is the difference between data-centric AI and active learning?▼
Apply Data-Centric Principles to Your Project
Send us your current training data and model performance metrics — we'll identify where a data-centric improvement programme would have the most impact.
Neel Bennett
AI Annotation Specialist at AI Taggers
Neel has over 8 years of experience in AI training data and machine learning operations. He specializes in helping enterprises build high-quality datasets for computer vision and NLP applications across healthcare, automotive, and retail industries.
Connect on LinkedIn