AuthorityAEO Guide

AI Hallucinations Start in the Data: The Annotation Link

Hallucination is framed as a model problem. The evidence says it is mostly a data problem — and the specific data that matters most is annotation quality.

4 October 202613 min read

Quick answer

AI hallucinations are confident, incorrect outputs generated by language models. They start in the data: mislabelled training examples, inconsistent annotations, and RLHF preference data that rewards fluency over accuracy each teach models that confident-sounding wrong answers are acceptable. Reducing hallucination rates requires systematic annotation quality controls — inter-annotator agreement measurement, gold-set calibration, and label error auditing — applied before training, not after deployment.

The Hallucination Problem Is Framed Wrong

When a language model states that a historical figure died in the wrong year, cites a paper that doesn't exist, or invents a legal precedent with perfect grammatical confidence, the immediate question asked is: “What's wrong with the model?” The more productive question is: “What did the training data teach it?”

Models do not hallucinate randomly. They hallucinate in patterns that reflect their training signal. A model that consistently invents plausible-sounding citations has been trained on examples where generating plausible-sounding text was the objective — either through label noise that rewarded incorrect but coherent outputs, or through RLHF preference annotation that consistently selected fluent responses over accurate ones.

The data dimension of hallucination has been systematically underweighted in public discourse, because model architecture is more legible and more publishable than annotation quality. This post addresses what the evidence actually shows about where hallucinations originate, and what annotation discipline is required to reduce them.

The Scale of the Label Noise Problem

A 2021 study by Northcutt, Jiang, and Chuang — “Pervasive Label Errors in Test Sets Destabilise Machine Learning Benchmarks” — audited label quality across 10 widely used benchmark datasets including ImageNet, Amazon Reviews, and the IMDB sentiment dataset. They found average label error rates of 3.3–6.0% across these datasets. Correcting these errors improved model performance on 9 of 10 benchmarks.

For production training sets — which are annotated under commercial conditions with mixed annotator pools, time pressure, and iterative guidelines — label error rates of 5–15% are common. A dataset with 10% label noise contains tens of thousands of examples explicitly teaching the model incorrect mappings between inputs and outputs.

The model does not know which examples are noisy. It treats every labelled example as signal. When a substantial fraction of that signal is incorrect, the model learns to produce incorrect outputs with the same confidence it produces correct ones — because in training, confidence was always rewarded regardless of accuracy.

Three Annotation Pathways to Hallucination

1. Label noise in supervised fine-tuning data

Supervised fine-tuning (SFT) data teaches a pre-trained language model how to respond to instructions. Every labelled example is a demonstration of the desired behaviour. When annotators incorrectly transcribe facts, misclassify entities, or mark incorrect answers as correct in QA datasets, the SFT process teaches the model that those incorrect outputs are the target.

The particularly dangerous case is instruction-following data where the “ideal response” has been written or selected by annotators who themselves had incorrect domain knowledge — a non-specialist annotator writing the “gold standard” answer to a medical or legal question, for example. The model is trained to reproduce that incorrect answer as authoritative.

2. Inconsistent annotation that teaches ambiguity

When annotators label the same or similar inputs inconsistently — because guidelines are ambiguous, calibration is inadequate, or the annotator pool is too mixed — the model learns that multiple different outputs are valid for the same input. During inference, it samples from that learned distribution and may produce whichever response it has been calibrated to generate most confidently, regardless of factual accuracy.

Inter-annotator agreement (IAA) scores below 0.7 Cohen's kappa on NLP tasks reliably predict downstream model inconsistency. The model's inconsistency is not a model architecture problem — it is a direct reflection of the annotation inconsistency it was trained on. Our guide to Cohen's kappa in annotation quality covers what these thresholds mean in practice.

3. RLHF preference data that rewards fluency over accuracy

Reinforcement learning from human feedback (RLHF) is where the hallucination problem becomes structurally entrenched if annotation is not disciplined. In RLHF, human annotators compare pairs of model responses and select the better one. Their preferences are used to train a reward model that shapes subsequent model behaviour.

Research from Anthropic, OpenAI, and academic groups has consistently found that human annotators without domain expertise tend to prefer responses that are well-structured, confidently phrased, and appropriately lengthy — over responses that are accurate but hedged, shorter, or that acknowledge uncertainty. When these preferences are encoded as reward signal, the model is literally being trained to be confidently wrong.

Is Your Training Data Hallucination-Prone?

A systematic data QA audit before training is far cheaper than post-deployment remediation. Our data QA and validation service identifies label errors, inconsistency patterns, and RLHF preference data problems before they reach your model.

Request a free dataset audit

Case Study: Reducing Hallucination in a Legal-Tech NLP Model

An Australian legal-tech company came to us with a contract clause extraction model that was producing hallucinated clause summaries at a rate that made the output unusable for legal review. The model was extracting correct clause boundaries but generating summaries that mixed elements from different clauses, invented obligations that weren't present, and omitted key conditions.

A systematic audit of the SFT dataset using confident learning (Cleanlab) identified a label error rate of 12.3% in the summary annotations. The primary error types were omission errors (annotators had left out qualifying conditions when writing gold-standard summaries) and false positive entity inclusions (annotators had included parties from adjacent clauses).

Our data QA and validation process involved: re-annotating the flagged 12.3% with a specialist legal annotator pool, implementing a dual-annotation plus adjudication workflow for all subsequent data, and establishing a gold-set calibration suite of 200 verified clauses used for annotator onboarding.

After retraining on the cleaned dataset:

The model architecture was unchanged. The only variable was annotation quality. This is the pattern we see repeatedly: hallucination problems that are attributed to model limitations are in practice annotation quality problems that respond to data remediation.

What Annotation-Grade Data Actually Requires

There is no universally agreed definition of “annotation-grade” data quality, but for teams concerned about hallucination, the minimum viable controls are:

Inter-annotator agreement measurement. Every annotation task should have IAA measured on a representative sample before full production begins. For classification tasks, Cohen's kappa above 0.7 is a common threshold. For more complex tasks (summarisation, preference ranking), task-specific rubrics with documented adjudication criteria are required. Without IAA measurement, you have no signal about whether your annotation guidelines are producing consistent labels.

Gold-set calibration. A gold set is a collection of examples with verified correct labels, used to test annotator accuracy before production work begins and periodically during production. Gold-set failure rates above 5–10% indicate a calibration problem that will propagate into the full dataset. Well-designed annotation guidelines are the prerequisite — our guide on writing annotation guidelines that don't need constant revision covers the structural requirements.

Domain-expert annotators for specialised tasks. For any task where factual accuracy is a quality dimension — medical, legal, financial, scientific — annotators without domain knowledge will introduce systematic factual errors that are invisible to standard IAA measurement. These errors directly cause hallucination. The solution is annotator credentialing, not platform features.

Systematic label error auditing before training. A pre-training label audit using confident learning or similar techniques can catch the high-noise subset of a dataset for targeted review. For large datasets, even auditing the top 5–10% of most likely erroneous examples and correcting them before training materially reduces downstream hallucination rates.

The RLHF Design Problem

RLHF preference annotation is the highest-stakes annotation task in contemporary AI development. The preferences collected by human annotators directly determine what behaviour the reward model incentivises — and therefore what the final model optimises for.

The design failure that most reliably produces hallucination-prone RLHF data is annotator briefing that emphasises helpfulness, coherence, and tone without explicit instruction to prefer accurate responses over confident-but-incorrect ones. Without explicit accuracy-preference guidelines, annotators default to preferring responses that sound authoritative — because sounding authoritative correlates with helpfulness in most of their lived experience.

Effective RLHF annotation guidelines for factual tasks require: explicit instruction to prefer responses that acknowledge uncertainty over responses that assert false certainty; domain-expert annotators for domain-specific comparisons; and preference-pair designs that force annotators to evaluate factual accuracy rather than just tone. Our guide on RLHF data collection for production LLMs covers these design requirements in detail.

A 2023 study from the University of Edinburgh found that RLHF models trained with accuracy-first preference guidelines showed a 23% reduction in factual error rate on TruthfulQA compared to models trained with standard helpfulness-first annotation — using the same base model and the same number of preference pairs. The annotation guidelines were the only variable.

Practical Steps for Teams Building Hallucination-Resistant Models

For ML teams currently building or retraining models with hallucination problems, the data-first intervention sequence is:

  1. Audit your SFT dataset for label error rate using confident learning. A rate above 8% warrants full-dataset review before retraining.
  2. Measure IAA retrospectively on a stratified sample of your existing annotations. Kappa below 0.65 on classification tasks indicates systemic guideline problems.
  3. Review your RLHF preference guidelines for accuracy-preference language. If accuracy is not explicitly prioritised over fluency, revise the guidelines and re-annotate a sample to verify the effect.
  4. Implement gold-set testing before any new annotation begins. The gold set should include examples where the correct answer requires domain knowledge, to catch non-expert annotator errors.
  5. Apply a QA relabeling process to the highest-error segments before retraining — see our guide on annotation QA and relabeling for the workflow.

The Bottom Line

AI hallucinations are a data quality problem more often than they are a model architecture problem. Label noise, annotation inconsistency, and RLHF preference guidelines that reward fluency over accuracy are each directly trainable pathways to confabulation. None of them require a better model to fix — they require better annotation.

The good news is that annotation quality is more directly controllable than model architecture. A systematic data QA and validation programme applied before training consistently produces larger reductions in hallucination rate than architectural interventions applied after training on poor-quality data. The legal-tech case study above is representative: the model did not change; the data did.

Frequently Asked Questions

Why do AI models hallucinate?▼
AI models hallucinate for several overlapping reasons, but the most directly addressable is training data quality. Mislabelled examples teach models that certain incorrect outputs are correct; annotation inconsistency teaches models that multiple different outputs are valid for the same input; and RLHF preference data that rewards fluency over accuracy directly trains models to produce confident-sounding but incorrect responses. Model architecture improvements help, but they cannot compensate for systematically noisy training data.
How does bad annotation cause AI hallucinations?▼
Bad annotation causes hallucinations by teaching the model incorrect signal. Label noise causes the model to learn wrong input-output mappings. Annotation inconsistency causes the model to generalise poorly, producing different answers for equivalent inputs. RLHF preference mislabels train the model to optimise for fluency rather than accuracy. Each of these is a data quality problem with a data quality solution: systematic IAA measurement, gold-set calibration, and pre-training label error auditing.
Can data QA actually reduce AI hallucination rates?▼
Yes. Research consistently shows that correcting label errors in training data improves model accuracy. In the legal-tech case study described in this post, cleaning a 12.3% label error rate reduced clause summary hallucination from 31.4% to 6.8% with no model architecture changes. For RLHF models, accuracy-first preference guidelines have been shown to reduce factual error rates by 23% compared to standard helpfulness-first annotation.
What types of annotation errors are most likely to cause hallucinations?▼
The three most hallucination-prone annotation error types are: RLHF preference mislabels (selecting fluent-but-incorrect over accurate-but-hedged responses); omission errors in SFT data (leaving out key facts or qualifications in gold-standard answers); and guideline inconsistency that produces conflicting labels for equivalent inputs. All three are preventable with proper annotation protocol design and QA controls.
How do you find annotation errors in an existing training dataset?▼
The most practical approaches for large datasets are: confident learning (using model predictions to flag likely label errors, implemented in the open-source Cleanlab library); retrospective IAA measurement on a stratified sample; and gold-set comparison. For RLHF data specifically, having domain experts review a sample of preference pairs for accuracy-vs-fluency trade-offs often reveals systematic annotator bias within a few hundred examples.
Is this a problem that only affects large language models?▼
No. The same mechanisms apply to any supervised model: computer vision models trained on mislabelled images, NER models trained on inconsistently annotated entities, and recommendation systems trained on noisy user feedback all exhibit similar patterns. Hallucination as a term is most associated with language models, but the underlying phenomenon — confident incorrect outputs attributable to label noise — is universal across ML task types.
Free Sample · 24-48 hours

Audit Your Training Data for Hallucination Risk

We'll analyse a sample of your annotation data, measure label error rate, and identify the fixes most likely to reduce hallucination in your next training run.

No commitment. NDA available on request. We respond within 24 hours, often the same day for Gulf-region inquiries.

Neel Bennett

AI Annotation Specialist at AI Taggers

Neel has over 8 years of experience in AI training data and machine learning operations. He specializes in helping enterprises build high-quality datasets for computer vision and NLP applications across healthcare, automotive, and retail industries.

Connect on LinkedIn