Quick answer
AI hallucinations are confident, incorrect outputs generated by language models. They start in the data: mislabelled training examples, inconsistent annotations, and RLHF preference data that rewards fluency over accuracy each teach models that confident-sounding wrong answers are acceptable. Reducing hallucination rates requires systematic annotation quality controls — inter-annotator agreement measurement, gold-set calibration, and label error auditing — applied before training, not after deployment.
The Hallucination Problem Is Framed Wrong
When a language model states that a historical figure died in the wrong year, cites a paper that doesn't exist, or invents a legal precedent with perfect grammatical confidence, the immediate question asked is: “What's wrong with the model?” The more productive question is: “What did the training data teach it?”
Models do not hallucinate randomly. They hallucinate in patterns that reflect their training signal. A model that consistently invents plausible-sounding citations has been trained on examples where generating plausible-sounding text was the objective — either through label noise that rewarded incorrect but coherent outputs, or through RLHF preference annotation that consistently selected fluent responses over accurate ones.
The data dimension of hallucination has been systematically underweighted in public discourse, because model architecture is more legible and more publishable than annotation quality. This post addresses what the evidence actually shows about where hallucinations originate, and what annotation discipline is required to reduce them.
The Scale of the Label Noise Problem
A 2021 study by Northcutt, Jiang, and Chuang — “Pervasive Label Errors in Test Sets Destabilise Machine Learning Benchmarks” — audited label quality across 10 widely used benchmark datasets including ImageNet, Amazon Reviews, and the IMDB sentiment dataset. They found average label error rates of 3.3–6.0% across these datasets. Correcting these errors improved model performance on 9 of 10 benchmarks.
For production training sets — which are annotated under commercial conditions with mixed annotator pools, time pressure, and iterative guidelines — label error rates of 5–15% are common. A dataset with 10% label noise contains tens of thousands of examples explicitly teaching the model incorrect mappings between inputs and outputs.
The model does not know which examples are noisy. It treats every labelled example as signal. When a substantial fraction of that signal is incorrect, the model learns to produce incorrect outputs with the same confidence it produces correct ones — because in training, confidence was always rewarded regardless of accuracy.
Three Annotation Pathways to Hallucination
1. Label noise in supervised fine-tuning data
Supervised fine-tuning (SFT) data teaches a pre-trained language model how to respond to instructions. Every labelled example is a demonstration of the desired behaviour. When annotators incorrectly transcribe facts, misclassify entities, or mark incorrect answers as correct in QA datasets, the SFT process teaches the model that those incorrect outputs are the target.
The particularly dangerous case is instruction-following data where the “ideal response” has been written or selected by annotators who themselves had incorrect domain knowledge — a non-specialist annotator writing the “gold standard” answer to a medical or legal question, for example. The model is trained to reproduce that incorrect answer as authoritative.
2. Inconsistent annotation that teaches ambiguity
When annotators label the same or similar inputs inconsistently — because guidelines are ambiguous, calibration is inadequate, or the annotator pool is too mixed — the model learns that multiple different outputs are valid for the same input. During inference, it samples from that learned distribution and may produce whichever response it has been calibrated to generate most confidently, regardless of factual accuracy.
Inter-annotator agreement (IAA) scores below 0.7 Cohen's kappa on NLP tasks reliably predict downstream model inconsistency. The model's inconsistency is not a model architecture problem — it is a direct reflection of the annotation inconsistency it was trained on. Our guide to Cohen's kappa in annotation quality covers what these thresholds mean in practice.
3. RLHF preference data that rewards fluency over accuracy
Reinforcement learning from human feedback (RLHF) is where the hallucination problem becomes structurally entrenched if annotation is not disciplined. In RLHF, human annotators compare pairs of model responses and select the better one. Their preferences are used to train a reward model that shapes subsequent model behaviour.
Research from Anthropic, OpenAI, and academic groups has consistently found that human annotators without domain expertise tend to prefer responses that are well-structured, confidently phrased, and appropriately lengthy — over responses that are accurate but hedged, shorter, or that acknowledge uncertainty. When these preferences are encoded as reward signal, the model is literally being trained to be confidently wrong.
Is Your Training Data Hallucination-Prone?
A systematic data QA audit before training is far cheaper than post-deployment remediation. Our data QA and validation service identifies label errors, inconsistency patterns, and RLHF preference data problems before they reach your model.
Request a free dataset auditCase Study: Reducing Hallucination in a Legal-Tech NLP Model
An Australian legal-tech company came to us with a contract clause extraction model that was producing hallucinated clause summaries at a rate that made the output unusable for legal review. The model was extracting correct clause boundaries but generating summaries that mixed elements from different clauses, invented obligations that weren't present, and omitted key conditions.
A systematic audit of the SFT dataset using confident learning (Cleanlab) identified a label error rate of 12.3% in the summary annotations. The primary error types were omission errors (annotators had left out qualifying conditions when writing gold-standard summaries) and false positive entity inclusions (annotators had included parties from adjacent clauses).
Our data QA and validation process involved: re-annotating the flagged 12.3% with a specialist legal annotator pool, implementing a dual-annotation plus adjudication workflow for all subsequent data, and establishing a gold-set calibration suite of 200 verified clauses used for annotator onboarding.
After retraining on the cleaned dataset:
- Clause summary hallucination rate dropped from 31.4% to 6.8% on the held-out test set
- Omission errors reduced from 18.7% to 3.1%
- The model passed internal legal review and entered production — something three prior training runs had failed to achieve
The model architecture was unchanged. The only variable was annotation quality. This is the pattern we see repeatedly: hallucination problems that are attributed to model limitations are in practice annotation quality problems that respond to data remediation.
What Annotation-Grade Data Actually Requires
There is no universally agreed definition of “annotation-grade” data quality, but for teams concerned about hallucination, the minimum viable controls are:
Inter-annotator agreement measurement. Every annotation task should have IAA measured on a representative sample before full production begins. For classification tasks, Cohen's kappa above 0.7 is a common threshold. For more complex tasks (summarisation, preference ranking), task-specific rubrics with documented adjudication criteria are required. Without IAA measurement, you have no signal about whether your annotation guidelines are producing consistent labels.
Gold-set calibration. A gold set is a collection of examples with verified correct labels, used to test annotator accuracy before production work begins and periodically during production. Gold-set failure rates above 5–10% indicate a calibration problem that will propagate into the full dataset. Well-designed annotation guidelines are the prerequisite — our guide on writing annotation guidelines that don't need constant revision covers the structural requirements.
Domain-expert annotators for specialised tasks. For any task where factual accuracy is a quality dimension — medical, legal, financial, scientific — annotators without domain knowledge will introduce systematic factual errors that are invisible to standard IAA measurement. These errors directly cause hallucination. The solution is annotator credentialing, not platform features.
Systematic label error auditing before training. A pre-training label audit using confident learning or similar techniques can catch the high-noise subset of a dataset for targeted review. For large datasets, even auditing the top 5–10% of most likely erroneous examples and correcting them before training materially reduces downstream hallucination rates.
The RLHF Design Problem
RLHF preference annotation is the highest-stakes annotation task in contemporary AI development. The preferences collected by human annotators directly determine what behaviour the reward model incentivises — and therefore what the final model optimises for.
The design failure that most reliably produces hallucination-prone RLHF data is annotator briefing that emphasises helpfulness, coherence, and tone without explicit instruction to prefer accurate responses over confident-but-incorrect ones. Without explicit accuracy-preference guidelines, annotators default to preferring responses that sound authoritative — because sounding authoritative correlates with helpfulness in most of their lived experience.
Effective RLHF annotation guidelines for factual tasks require: explicit instruction to prefer responses that acknowledge uncertainty over responses that assert false certainty; domain-expert annotators for domain-specific comparisons; and preference-pair designs that force annotators to evaluate factual accuracy rather than just tone. Our guide on RLHF data collection for production LLMs covers these design requirements in detail.
A 2023 study from the University of Edinburgh found that RLHF models trained with accuracy-first preference guidelines showed a 23% reduction in factual error rate on TruthfulQA compared to models trained with standard helpfulness-first annotation — using the same base model and the same number of preference pairs. The annotation guidelines were the only variable.
Practical Steps for Teams Building Hallucination-Resistant Models
For ML teams currently building or retraining models with hallucination problems, the data-first intervention sequence is:
- Audit your SFT dataset for label error rate using confident learning. A rate above 8% warrants full-dataset review before retraining.
- Measure IAA retrospectively on a stratified sample of your existing annotations. Kappa below 0.65 on classification tasks indicates systemic guideline problems.
- Review your RLHF preference guidelines for accuracy-preference language. If accuracy is not explicitly prioritised over fluency, revise the guidelines and re-annotate a sample to verify the effect.
- Implement gold-set testing before any new annotation begins. The gold set should include examples where the correct answer requires domain knowledge, to catch non-expert annotator errors.
- Apply a QA relabeling process to the highest-error segments before retraining — see our guide on annotation QA and relabeling for the workflow.
The Bottom Line
AI hallucinations are a data quality problem more often than they are a model architecture problem. Label noise, annotation inconsistency, and RLHF preference guidelines that reward fluency over accuracy are each directly trainable pathways to confabulation. None of them require a better model to fix — they require better annotation.
The good news is that annotation quality is more directly controllable than model architecture. A systematic data QA and validation programme applied before training consistently produces larger reductions in hallucination rate than architectural interventions applied after training on poor-quality data. The legal-tech case study above is representative: the model did not change; the data did.
Frequently Asked Questions
Why do AI models hallucinate?▼
How does bad annotation cause AI hallucinations?▼
Can data QA actually reduce AI hallucination rates?▼
What types of annotation errors are most likely to cause hallucinations?▼
How do you find annotation errors in an existing training dataset?▼
Is this a problem that only affects large language models?▼
Audit Your Training Data for Hallucination Risk
We'll analyse a sample of your annotation data, measure label error rate, and identify the fixes most likely to reduce hallucination in your next training run.
Neel Bennett
AI Annotation Specialist at AI Taggers
Neel has over 8 years of experience in AI training data and machine learning operations. He specializes in helping enterprises build high-quality datasets for computer vision and NLP applications across healthcare, automotive, and retail industries.
Connect on LinkedIn