Quick answer
Medical AI misses tumours because of training data failures, not just model failures. When histopathology or imaging training datasets contain systematic annotation errors — labels applied by insufficiently credentialed annotators, inconsistent boundary definitions, or underrepresented rare presentations — the model learns the error, not the pathology. Annotation provenance — documented, traceable, qualified-annotator records for every label — is the mechanism that prevents this. For FDA-regulated AI as a Software as a Medical Device (SaMD), it is also a regulatory requirement under 21 CFR Part 11.
The Scenario: How a Missed Finding Traces Back to Annotation
A digital pathology company develops a histopathology AI system to screen colorectal biopsy slides for adenocarcinoma. The training dataset contains 42,000 annotated whole-slide image (WSI) patches, assembled from two academic hospital datasets and one commercial source. Validation shows sensitivity of 91.4% on the holdout set. The company proceeds to a pilot clinical deployment.
Six months post-deployment, a clinical audit reveals that the model is performing significantly worse on samples from one site — sensitivity on adenocarcinoma drops to 73.2% for that cohort. A subset of false negatives is pulled for review. In 14 of the 19 cases, the tumour region is present and visible but not flagged.
The post-mortem begins. Data scientists compare the affected cases to the training distribution. Radiologists review the false-negative patches. The finding: the affected cases predominantly show a moderately differentiated adenocarcinoma pattern — a histological variant that is diagnostic but can appear less dramatically abnormal than the high-grade variants more common in the training data.
The training dataset contained 3,800 patches labelled as "adenocarcinoma" — but only 280 of those were moderately differentiated. Of those 280, 61 had been labelled by a pathology resident without attending pathologist review, and retrospective review found that 23 of those 61 labels were either incorrectly applied or had ambiguous boundary definitions. The model had a statistically insufficient, partially mislabelled training set for the exact presentation it was most likely to miss in the clinic.
Why Medical AI Errors Are Data Errors
The framing of medical AI failures as "model errors" is persistent and largely wrong. A 2022 study published in Nature Medicine analysed 14 publicly available medical imaging datasets and found that label noise correlated directly with model sensitivity loss on minority subgroup presentations. The paper concluded that in 9 of the 14 datasets, systematic label errors were a more significant contributor to subgroup performance gaps than architectural or data quantity factors.
A 2021 analysis in The Lancet Digital Health examined label error rates across 10 widely used clinical imaging benchmarks and found rates ranging from 3.3% to 12.1%. At 5%, a training set of 40,000 images contains 2,000 mislabelled examples. For a rare finding type represented by only 800 positive examples in the dataset, a 5% error rate in that class means 40 of 800 positives are mislabelled — a 5% label error rate becomes a potential 15–20 percentage point sensitivity reduction for that finding.
The failure mode is not random noise, which can sometimes be survived with enough data volume. It is systematic bias: annotation errors that cluster around specific histological patterns, imaging conditions, or demographic subgroups — exactly the presentations the model needs to handle correctly in clinical use.
What Annotation Provenance Actually Means
Annotation provenance is the documented chain of records that traces each label back to: the qualified annotator who created it; the date and version of annotation guidelines in force at the time; the review or adjudication process applied; and any subsequent revision with its rationale. It is not a nice-to-have for medical AI — for FDA-regulated AI as a Software as a Medical Device, it is a requirement under 21 CFR Part 11 for electronic records.
In practice, most academic and early-stage medical imaging datasets lack full provenance. Labels may be attributed to an institution rather than a named qualified individual. Revisions may overwrite originals without logging. Adjudication outcomes may not be recorded. When the FDA reviews a SaMD submission, these gaps become regulatory blockers.
Clinical-grade provenance requires: unique annotator IDs with documented credentials (board certification level, years of subspecialty experience), timestamped audit trails for every label and revision, access controls preventing unauthorised modification, and exportable provenance records in a format compatible with FDA submission packages. Annotation platforms used for regulated medical AI must support all of these — and the annotation vendor must be able to certify compliance.
Need clinical-grade histopathology annotation?
AI Taggers provides histopathology annotation with board-certified pathologist annotators, multi-pathologist adjudication, and FDA 21 CFR Part 11-aligned provenance documentation.
Discuss your datasetCase Study: Recovering Sensitivity Through Annotation Remediation
A medical AI company developing a lung adenocarcinoma detection system for CT scans identified a systematic false-negative pattern six months before a planned regulatory submission. Their model showed overall sensitivity of 87.3% but dropped to 71.6% for ground-glass opacity (GGO)-predominant adenocarcinoma — a subtype that accounts for approximately 30% of clinical presentation in their target population.
The diagnosis: A provenance audit of the training dataset revealed that GGO-predominant cases had been labelled by radiologists without thoracic imaging subspecialty — adequate for consolidative adenocarcinoma but insufficiently trained on the subtle visual features of GGO patterns. Inter-annotator agreement (Cohen's kappa) for GGO cases was 0.61, compared with 0.89 for consolidative cases. The lower agreement reflected genuine annotator uncertainty, which translated into label noise.
The remediation project: AI Taggers coordinated a targeted remediation of 1,840 GGO-predominant patches. Two board-certified thoracic radiologists annotated each patch independently. Cases with disagreement (28% of the remediation set) underwent adjudication by a senior thoracic radiologist with subspecialty lung oncology experience. All annotations were recorded with full 21 CFR Part 11-compliant provenance documentation.
The results: After retraining on the remediated dataset:
- Sensitivity for GGO-predominant adenocarcinoma improved from 71.6% to 88.4%
- Overall model sensitivity increased from 87.3% to 91.7%
- False negative rate for GGO presentations fell by 59%
- The provenance documentation package met FDA pre-submission review requirements without supplementary requests
The annotation remediation project took nine weeks from scoping to delivery. The alternative — discovering the GGO performance gap post-submission or post-deployment — would have meant regulatory delay, retraining cost, and potential clinical harm.
The Five Annotation Disciplines That Prevent Medical AI Failures
1. Credentialed annotators, not generic medical professionals
Histopathology annotation requires board-certified pathologists — not general practitioners, nurses, or even physicians from other specialties. For subspecialty tasks (neuropathology, haematopathology, dermatopathology), subspecialty credentials matter. Annotator credential documentation is not optional for clinical-grade datasets; it is a provenance requirement.
2. Multi-annotator adjudication for ambiguous cases
Single-annotator labels on challenging histological cases are insufficient for clinical-grade training data. Cases where two independent annotators disagree should undergo formal adjudication — a senior or subspecialty expert reviews both labels and makes a documented determination. The adjudication record, including the disagreement and the rationale for the final label, forms part of the dataset provenance.
3. Inter-annotator agreement tracking by case type
IAA metrics like Cohen's kappa should be tracked not just overall but broken down by finding type, rarity, and imaging presentation. An overall kappa of 0.85 can mask a kappa of 0.61 for the rare or ambiguous cases that are most clinically important. Low kappa on a specific finding type is a signal to review annotation guidelines, increase adjudication, or bring in additional annotator expertise before those cases contaminate the training set.
4. Rare and edge-case representation
Medical AI models are tested in the clinic on the full distribution of presentations — including the rare ones that do not appear frequently in training datasets assembled from academic hospitals. Active sampling of rare finding types, underrepresented demographics, and atypical presentations is a dataset design requirement, not a nice-to-have. Without it, the model is being trained on the easy cases and evaluated on the hard ones.
5. FDA-aligned provenance from day one
Retrofitting provenance documentation onto an existing dataset is expensive and sometimes impossible if annotation records were not captured at the time. Building 21 CFR Part 11-compliant provenance into the annotation workflow from the start — annotator IDs, timestamps, audit trails, access controls — ensures that the documentation package for FDA submission reflects the actual dataset, not a reconstruction. AI Taggers' histopathology annotation service is structured around FDA-aligned provenance by default.
What to Do if You Suspect a Training Data Problem
If your medical AI model shows unexplained performance gaps on specific case types, demographics, or imaging conditions, the first step is a targeted provenance audit — not a model retraining run. Review the annotation records for the affected case types: who labelled them, under what guidelines, with what IAA, and whether they underwent adjudication.
Statistical tests for label noise — such as confident learning (Northcutt et al., 2021) or out-of-distribution detection on label distributions — can identify the cases most likely to be mislabelled before a clinician review. This narrows the expert review workload to the cases that matter most.
Related reading: histopathology whole-slide image annotation workflows, FDA 21 CFR Part 11 annotation documentation requirements, and inter-rater reliability in medical annotation.
Frequently Asked Questions
Why does medical AI miss tumours?
What is annotation provenance in medical AI?
Who should annotate histopathology images?
How common are annotation errors in medical AI datasets?
What does FDA 21 CFR Part 11 require for medical AI annotation?
Need Clinical-Grade Annotation with Full Provenance?
Tell us about your medical AI dataset and we'll scope a board-certified annotation project with FDA-aligned documentation.
Neel Bennett
AI Annotation Specialist at AI Taggers
Neel has over 8 years of experience in AI training data and machine learning operations. He specializes in helping enterprises build high-quality datasets for computer vision and NLP applications across healthcare, automotive, and retail industries.
Connect on LinkedIn