AuthorityThought Leadership

Data-Centric AI: Why the Model Stopped Being the Bottleneck

For most real-world AI tasks, improving data quality produces larger accuracy gains than improving model architecture. This is not a hypothesis — it has been demonstrated empirically across manufacturing, medical imaging, NLP, and computer vision. Data-centric AI is the name for treating this observation as an engineering discipline.

30 September 202613 min read

Direct answer

Data-centric AI is an approach to machine learning that prioritises systematic improvement of training data quality over model architecture iteration. The core evidence: Andrew Ng's 2021 experiments showed that consistent data quality improvement outperformed architecture switching on a speech recognition task, producing twice the accuracy gain at lower cost. For most production AI tasks with datasets under 100k examples, this pattern holds — the model is not the bottleneck; the data is.

The Shift That Changed How Production AI Teams Operate

For the first decade of the deep learning era, the dominant frame was model-centric: hold the dataset fixed and iterate on the model. The ImageNet competition rewarded architecture innovation. Benchmarks were fixed datasets. The field's metrics of progress were model accuracy on standard held-out sets, not data quality.

This made sense for academic research where the benchmark is the goal. It does not make sense for production AI, where the goal is a deployed system that performs correctly on real-world inputs. In production, the dataset is not fixed — it is collected, annotated, and continuously updated. And in practice, the annotation process is rarely managed with the same engineering rigour as the model training process.

The result: enterprise AI teams spending months iterating on model architecture while their training data contains systematic label errors, coverage gaps, and distribution mismatches that no model can compensate for. Data-centric AI names this failure mode and provides a corrective framework.

The Evidence: What Ng's Experiments Actually Showed

In 2021, Andrew Ng ran a structured experiment to compare model-centric and data-centric improvement strategies on a speech recognition task. The setup: a fixed baseline model, a fixed test set, and two improvement paths — one that iterated on model architecture and hyperparameters while holding data constant, one that systematically improved data quality while holding the model constant.

The data-centric path achieved approximately 2× the accuracy improvement of the model-centric path. The specific intervention was consistent labelling: ensuring that similar utterances were labelled consistently across the training set, that noise and artefacts were handled uniformly, and that annotation guidelines produced the same decisions for similar inputs. No new data was collected — existing data was relabelled more carefully.

This experiment was influential not because the result was surprising in isolation — practitioners had long observed that data quality mattered — but because it was structured, comparable, and public. It gave the data quality observation the empirical backing needed to change how teams allocate engineering effort.

Data improvement vs model improvement: typical outcomes

Model-centric approach
+3–6 pp
Typical accuracy gain from architecture switching on a production task with clean data
Data-centric approach
+8–20 pp
Typical accuracy gain from systematic annotation quality improvement on the same task

Source: Internal project benchmarks across 30+ enterprise annotation projects, 2024–2026. Ranges vary by task, dataset size, and baseline data quality.

What Data-Centric AI Means in Practice

Data-centric AI is not a specific algorithm or framework — it is a posture: a set of priorities and practices that treat data improvement as the primary lever for model improvement. In practice, it involves three interlocking activities:

1. Systematic data quality measurement

You cannot improve what you cannot measure. Data-centric AI programmes begin with a structured quality audit: measuring inter-annotator agreement (IAA) to quantify labelling consistency, constructing gold-standard sets to measure annotator accuracy, and profiling error rates by label class and data subtype to identify where quality is degrading model performance.

The Cohen's kappa coefficient is the standard IAA metric for categorical annotation tasks. A kappa below 0.7 indicates a labelling consistency problem at the guidelines level. A kappa above 0.85 indicates the annotation task is well-defined and consistently executed. These thresholds give teams an objective signal for when to invest in annotation quality improvement versus when to move to model iteration.

2. Targeted data improvement, not blanket re-annotation

A common misreading of data-centric AI is that it means re-annotating everything. It does not. The most efficient data-centric programmes identify the specific slices of the training data that are most affecting model performance — a particular label class, a specific data source, a distribution gap in the training set — and direct improvement effort there.

This is where active learning and human-in-the-loop workflows become powerful: the model itself surfaces the examples it is most uncertain about, directing annotator attention to the data that will have the highest marginal impact on accuracy. When the active learning loop is correctly implemented, each annotation cycle produces compounding quality improvements rather than linear ones.

AI Taggers' custom annotation workflows are designed to integrate this kind of targeted improvement — starting from a data audit and prioritising relabelling effort based on measured impact, not uniform coverage.

3. Annotation infrastructure that supports iteration

Data-centric AI is iterative. Each improvement cycle produces a new training set, a new model, and new performance measurements that feed the next audit. This requires annotation infrastructure that supports version control of datasets, reproducible measurement of IAA, and systematic tracking of which label schema versions produced which model performance.

Many organisations do not have this infrastructure. They annotate once, train once, and discover problems in production. Data-centric AI requires a different operational model: annotation as an ongoing engineering activity, not a one-time procurement.

Ready to take a data-centric approach to your AI project?

AI Taggers designs custom annotation workflows built around iterative data quality improvement — from initial audit to production-grade pipeline.

See our custom annotation services

A Real Data-Centric Programme: NLP Classifier from 71% to 94% F1

An Australian insurtech company engaged AI Taggers to improve a claims-routing NLP classifier that had plateaued at 71% F1 on a 12-class routing taxonomy. Their ML team had run eight model iterations over six months — BERT, RoBERTa, domain-adapted variants — without breaking 75%. The engineering lead suspected a data problem but had no measurement infrastructure to confirm it.

The data-centric programme proceeded in three phases:

1

Audit phase (2 weeks)

Gold-set audit on 800 training records revealed a 14.7% label error rate overall, concentrated in three routing categories where the taxonomy had changed mid-project and old labels had not been updated. IAA measurement on a 200-record sample returned kappa of 0.61 — well below acceptable threshold. Two annotation guidelines documents were in use simultaneously, creating systematic annotation divergence.

14.7% label error ratekappa 0.61conflicting guidelines
2

Remediation phase (4 weeks)

Consolidated guidelines with 25 concrete decision-boundary examples per high-confusion class pair. Relabelled 6,200 records in the three high-error categories. Removed 340 records where claims text had been partially redacted, making label assignment ambiguous — these had been contaminating training since project inception. Re-measured IAA: kappa 0.86.

kappa 0.61 → 0.866,200 records relabelled340 contaminated records removed
3

Expansion phase (3 weeks)

Active learning run on the updated model surfaced 1,400 low-confidence examples concentrated in two routing classes with thin training coverage. Annotated these at double density (three annotators, adjudication on disagreements). Final retrain on same RoBERTa architecture used for the original baseline.

1,400 active-learning targeted recordssame model architecture

Results: same model, systematically improved data

71% → 94%
Macro F1
14.7% → 1.9%
Label error rate
9 weeks
Total programme duration
0
Architecture changes

The model that had plateaued at 71% for six months of architecture iteration reached 94% in nine weeks of data improvement. The architecture did not change. The data did. This is data-centric AI operating as intended.

Where Data-Centric AI Has the Most Impact

Data-centric approaches produce the largest gains in specific contexts. Understanding where the lever is largest helps teams prioritise investment:

Small to medium datasets (under 100k examples)

At small dataset scale, label errors represent a higher fraction of the total signal and the model has less opportunity to average out noise. Each mislabelled example has disproportionate influence. Data-centric improvement here is high-leverage.

Complex or ambiguous label schemas

When the annotation task requires nuanced judgement — subtle sentiment, overlapping entity types, multi-label classification with correlated categories — annotation consistency is intrinsically harder to achieve and intrinsically more impactful when achieved. Tasks like claims routing, medical coding, and moderation have this character.

Production systems that have drifted from their training distribution

Models trained 12–24 months ago on historical data increasingly diverge from the current production distribution. Expanding training data coverage to match current production inputs — a data-centric intervention — can recover performance without retraining from scratch.

Foundation model fine-tuning

Fine-tuning a large pre-trained model is almost entirely a data-centric problem. The base model has broad capability; the question is what examples to fine-tune it on. Quality of the fine-tuning dataset — typically hundreds to thousands of examples — dominates quality of the fine-tuned model.

High-stakes domains: medical, financial, legal

When annotation requires domain expertise — radiologists labelling pathology, clinicians annotating clinical text, compliance specialists reviewing financial documents — annotator quality variation is large and the cost of systematic errors is high. Data-centric principles, applied to expert annotation pipelines, produce large quality improvements at these quality tiers.

Data-Centric AI and Custom Annotation Workflows

The operational implication of data-centric AI is that annotation cannot be a commodity procurement: buy labels, run training, deploy model. It requires custom annotation workflows that are designed around the specific quality requirements of each project — with measurement built in from the start, not audited at the end.

A data-centric annotation workflow typically includes:

This infrastructure is not optional overhead — it is what makes the data-centric improvement loop possible. Without it, each annotation cycle is disconnected from the previous one, and the compounding quality improvements that data-centric AI promises are not achievable.

Applying Data-Centric Principles to Build vs Buy Decisions

Data-centric AI has direct implications for the build vs buy annotation decision. In-house annotation teams are often constructed around the model-centric assumption: label a dataset once, use it, move on. The data-centric model requires ongoing annotation capability — an audit-remediate-retrain cycle that does not have a natural end point.

For most organisations, sustaining the data-centric loop in-house is feasible for a single high-priority project but difficult across a portfolio of AI initiatives. External annotation partners who have invested in measurement infrastructure, expert annotator pools, and IAA-based QA protocols can run the improvement cycle at lower cost than in-house teams while producing higher-quality data.

The evaluation criterion shifts accordingly: the right annotation partner for a data-centric programme is not necessarily the cheapest per label, but the one whose pipeline produces the highest measurable data quality and whose reporting infrastructure makes the improvement cycle visible. For a transparent breakdown of what production-quality annotation costs by task type, see our 2026 annotation pricing guide.

Data-Centric AI in 2026: Where the Field Has Settled

Five years after Ng's initial 2021 campaign, data-centric AI has moved from a corrective framing to a mainstream practice in enterprise ML. The major shifts:

Foundation model fine-tuning has made data quality even more central. When the base model is pre-trained at massive scale, the differentiation between products built on it is almost entirely in fine-tuning data quality. RLHF and DPO training for alignment — preference data annotation — is a canonical data-centric activity where annotation quality directly determines whether the fine-tuned model is helpful or harmful.

Tooling for data quality measurement has matured. CleanLab and Confident Learning from MIT formalizsed the algorithmic detection of label errors at scale. Label Studio, Argilla, and enterprise platforms now include IAA measurement as standard features. The measurement infrastructure that data-centric AI requires is no longer bespoke engineering.

The "synthetic data will replace annotation" thesis has been tested and found partial. Synthetic data is effective for specific tasks — object detection augmentation, test set expansion, edge-case generation — but systematic tests have shown that synthetic data for NLP fine-tuning consistently underperforms real human-annotated data when the task requires genuine language understanding. The synthetic vs annotated data question is now well-characterised: complementary rather than substitutable for most production tasks.

Frequently Asked Questions

What is data-centric AI in simple terms?▼
Data-centric AI is the practice of treating your training data as the primary thing to improve, rather than your model. Instead of trying different model architectures to get better AI performance, you focus on making your training data more accurate, more consistent, and more representative of what the model will encounter in production. The evidence shows that for most real-world tasks, this produces larger accuracy gains than model iteration.
Is data-centric AI the same as data quality?▼
Data quality is a component of data-centric AI, but data-centric AI is broader. It includes not just quality measurement and improvement but also data coverage (ensuring the training set represents the production distribution), annotation consistency (ensuring similar inputs get the same labels), and the operational infrastructure to run improvement cycles systematically. Data quality is the goal; data-centric AI is the programme for achieving it.
How do you know when to focus on data vs model?▼
The fastest diagnostic is to measure inter-annotator agreement on your training data. If agreement (kappa) is below 0.75, focus on data first — the labelling guidelines are too ambiguous and any model will learn the noise. If agreement is strong but performance is still poor, check whether your training distribution matches production. If both look sound, the problem is more likely architectural or representational.
Can data-centric AI work for LLM fine-tuning?▼
Yes — in fact, fine-tuning is one of the contexts where data-centric principles matter most. The base model already has broad capability; the fine-tuning data determines what specific capabilities and behaviours it learns. Quality of the fine-tuning dataset — in terms of consistency, coverage, and alignment with intended behaviour — dominates quality of the fine-tuned model in most experiments.
What is the difference between data-centric AI and active learning?▼
Data-centric AI is the overarching approach that treats data quality as the primary lever for model improvement. Active learning is one specific technique within that approach: using model uncertainty signals to direct annotation effort toward the examples with the highest marginal value. Active learning accelerates data-centric improvement cycles but is not the same thing as data-centric AI.
Free Sample · 24-48 hours

Apply Data-Centric Principles to Your Project

Send us your current training data and model performance metrics — we'll identify where a data-centric improvement programme would have the most impact.

No commitment. NDA available on request. We respond within 24 hours, often the same day for Gulf-region inquiries.

Neel Bennett

AI Annotation Specialist at AI Taggers

Neel has over 8 years of experience in AI training data and machine learning operations. He specializes in helping enterprises build high-quality datasets for computer vision and NLP applications across healthcare, automotive, and retail industries.

Connect on LinkedIn