MedicalAuthority

When Medical AI Misses a Tumour: A Data-Provenance Post-Mortem

Medical AI failures almost always have a data provenance problem at their root. Here is how annotation errors cause diagnostic AI to miss critical findings — and what rigorous provenance disciplines prevent.

8 October 202613 min read

Quick answer

Medical AI misses tumours because of training data failures, not just model failures. When histopathology or imaging training datasets contain systematic annotation errors — labels applied by insufficiently credentialed annotators, inconsistent boundary definitions, or underrepresented rare presentations — the model learns the error, not the pathology. Annotation provenance — documented, traceable, qualified-annotator records for every label — is the mechanism that prevents this. For FDA-regulated AI as a Software as a Medical Device (SaMD), it is also a regulatory requirement under 21 CFR Part 11.

The Scenario: How a Missed Finding Traces Back to Annotation

A digital pathology company develops a histopathology AI system to screen colorectal biopsy slides for adenocarcinoma. The training dataset contains 42,000 annotated whole-slide image (WSI) patches, assembled from two academic hospital datasets and one commercial source. Validation shows sensitivity of 91.4% on the holdout set. The company proceeds to a pilot clinical deployment.

Six months post-deployment, a clinical audit reveals that the model is performing significantly worse on samples from one site — sensitivity on adenocarcinoma drops to 73.2% for that cohort. A subset of false negatives is pulled for review. In 14 of the 19 cases, the tumour region is present and visible but not flagged.

The post-mortem begins. Data scientists compare the affected cases to the training distribution. Radiologists review the false-negative patches. The finding: the affected cases predominantly show a moderately differentiated adenocarcinoma pattern — a histological variant that is diagnostic but can appear less dramatically abnormal than the high-grade variants more common in the training data.

The training dataset contained 3,800 patches labelled as "adenocarcinoma" — but only 280 of those were moderately differentiated. Of those 280, 61 had been labelled by a pathology resident without attending pathologist review, and retrospective review found that 23 of those 61 labels were either incorrectly applied or had ambiguous boundary definitions. The model had a statistically insufficient, partially mislabelled training set for the exact presentation it was most likely to miss in the clinic.

Why Medical AI Errors Are Data Errors

The framing of medical AI failures as "model errors" is persistent and largely wrong. A 2022 study published in Nature Medicine analysed 14 publicly available medical imaging datasets and found that label noise correlated directly with model sensitivity loss on minority subgroup presentations. The paper concluded that in 9 of the 14 datasets, systematic label errors were a more significant contributor to subgroup performance gaps than architectural or data quantity factors.

A 2021 analysis in The Lancet Digital Health examined label error rates across 10 widely used clinical imaging benchmarks and found rates ranging from 3.3% to 12.1%. At 5%, a training set of 40,000 images contains 2,000 mislabelled examples. For a rare finding type represented by only 800 positive examples in the dataset, a 5% error rate in that class means 40 of 800 positives are mislabelled — a 5% label error rate becomes a potential 15–20 percentage point sensitivity reduction for that finding.

The failure mode is not random noise, which can sometimes be survived with enough data volume. It is systematic bias: annotation errors that cluster around specific histological patterns, imaging conditions, or demographic subgroups — exactly the presentations the model needs to handle correctly in clinical use.

What Annotation Provenance Actually Means

Annotation provenance is the documented chain of records that traces each label back to: the qualified annotator who created it; the date and version of annotation guidelines in force at the time; the review or adjudication process applied; and any subsequent revision with its rationale. It is not a nice-to-have for medical AI — for FDA-regulated AI as a Software as a Medical Device, it is a requirement under 21 CFR Part 11 for electronic records.

In practice, most academic and early-stage medical imaging datasets lack full provenance. Labels may be attributed to an institution rather than a named qualified individual. Revisions may overwrite originals without logging. Adjudication outcomes may not be recorded. When the FDA reviews a SaMD submission, these gaps become regulatory blockers.

Clinical-grade provenance requires: unique annotator IDs with documented credentials (board certification level, years of subspecialty experience), timestamped audit trails for every label and revision, access controls preventing unauthorised modification, and exportable provenance records in a format compatible with FDA submission packages. Annotation platforms used for regulated medical AI must support all of these — and the annotation vendor must be able to certify compliance.

Need clinical-grade histopathology annotation?

AI Taggers provides histopathology annotation with board-certified pathologist annotators, multi-pathologist adjudication, and FDA 21 CFR Part 11-aligned provenance documentation.

Discuss your dataset

Case Study: Recovering Sensitivity Through Annotation Remediation

A medical AI company developing a lung adenocarcinoma detection system for CT scans identified a systematic false-negative pattern six months before a planned regulatory submission. Their model showed overall sensitivity of 87.3% but dropped to 71.6% for ground-glass opacity (GGO)-predominant adenocarcinoma — a subtype that accounts for approximately 30% of clinical presentation in their target population.

The diagnosis: A provenance audit of the training dataset revealed that GGO-predominant cases had been labelled by radiologists without thoracic imaging subspecialty — adequate for consolidative adenocarcinoma but insufficiently trained on the subtle visual features of GGO patterns. Inter-annotator agreement (Cohen's kappa) for GGO cases was 0.61, compared with 0.89 for consolidative cases. The lower agreement reflected genuine annotator uncertainty, which translated into label noise.

The remediation project: AI Taggers coordinated a targeted remediation of 1,840 GGO-predominant patches. Two board-certified thoracic radiologists annotated each patch independently. Cases with disagreement (28% of the remediation set) underwent adjudication by a senior thoracic radiologist with subspecialty lung oncology experience. All annotations were recorded with full 21 CFR Part 11-compliant provenance documentation.

The results: After retraining on the remediated dataset:

The annotation remediation project took nine weeks from scoping to delivery. The alternative — discovering the GGO performance gap post-submission or post-deployment — would have meant regulatory delay, retraining cost, and potential clinical harm.

The Five Annotation Disciplines That Prevent Medical AI Failures

1. Credentialed annotators, not generic medical professionals

Histopathology annotation requires board-certified pathologists — not general practitioners, nurses, or even physicians from other specialties. For subspecialty tasks (neuropathology, haematopathology, dermatopathology), subspecialty credentials matter. Annotator credential documentation is not optional for clinical-grade datasets; it is a provenance requirement.

2. Multi-annotator adjudication for ambiguous cases

Single-annotator labels on challenging histological cases are insufficient for clinical-grade training data. Cases where two independent annotators disagree should undergo formal adjudication — a senior or subspecialty expert reviews both labels and makes a documented determination. The adjudication record, including the disagreement and the rationale for the final label, forms part of the dataset provenance.

3. Inter-annotator agreement tracking by case type

IAA metrics like Cohen's kappa should be tracked not just overall but broken down by finding type, rarity, and imaging presentation. An overall kappa of 0.85 can mask a kappa of 0.61 for the rare or ambiguous cases that are most clinically important. Low kappa on a specific finding type is a signal to review annotation guidelines, increase adjudication, or bring in additional annotator expertise before those cases contaminate the training set.

4. Rare and edge-case representation

Medical AI models are tested in the clinic on the full distribution of presentations — including the rare ones that do not appear frequently in training datasets assembled from academic hospitals. Active sampling of rare finding types, underrepresented demographics, and atypical presentations is a dataset design requirement, not a nice-to-have. Without it, the model is being trained on the easy cases and evaluated on the hard ones.

5. FDA-aligned provenance from day one

Retrofitting provenance documentation onto an existing dataset is expensive and sometimes impossible if annotation records were not captured at the time. Building 21 CFR Part 11-compliant provenance into the annotation workflow from the start — annotator IDs, timestamps, audit trails, access controls — ensures that the documentation package for FDA submission reflects the actual dataset, not a reconstruction. AI Taggers' histopathology annotation service is structured around FDA-aligned provenance by default.

What to Do if You Suspect a Training Data Problem

If your medical AI model shows unexplained performance gaps on specific case types, demographics, or imaging conditions, the first step is a targeted provenance audit — not a model retraining run. Review the annotation records for the affected case types: who labelled them, under what guidelines, with what IAA, and whether they underwent adjudication.

Statistical tests for label noise — such as confident learning (Northcutt et al., 2021) or out-of-distribution detection on label distributions — can identify the cases most likely to be mislabelled before a clinician review. This narrows the expert review workload to the cases that matter most.

Related reading: histopathology whole-slide image annotation workflows, FDA 21 CFR Part 11 annotation documentation requirements, and inter-rater reliability in medical annotation.

Frequently Asked Questions

Why does medical AI miss tumours?
Medical AI systems miss tumours primarily due to training data failures: annotation errors by insufficiently credentialed annotators, inconsistent label boundaries across slides or frames, underrepresentation of rare or early-stage tumour presentations, and absence of edge-case examples from underrepresented demographic groups. A 2022 study in Nature Medicine found that label noise in medical imaging datasets correlates directly with model sensitivity loss on minority subgroup presentations.
What is annotation provenance in medical AI?
Annotation provenance is the documented chain of records tracing each label back to the qualified annotator who created it, the review process it underwent, the version of annotation guidelines in force at the time, and any adjudication decisions applied. For FDA-regulated AI/ML as a Software as a Medical Device (SaMD), provenance documentation is a regulatory requirement under 21 CFR Part 11.
Who should annotate histopathology images?
Histopathology images should be annotated by board-certified pathologists or, for clearly defined sub-tasks, by trained pathology residents under pathologist review. For challenging cases — rare tumour types, borderline lesions, multi-focal patterns — multi-pathologist adjudication is the standard for clinical-grade datasets.
How common are annotation errors in medical AI datasets?
Research published in The Lancet Digital Health (2021) found that publicly available medical imaging datasets contain label error rates of 3.3% to 12.1%. At 5%, a training set of 40,000 images contains 2,000 mislabelled examples — enough to materially degrade model sensitivity on affected finding types.
What does FDA 21 CFR Part 11 require for medical AI annotation?
FDA 21 CFR Part 11 requires electronic records and electronic signatures to be trustworthy, reliable, and equivalent to paper records. For medical AI annotation, this means unique annotator IDs with documented qualifications, timestamped audit trails of every label and revision, access controls, and backup/recovery procedures.
Free Sample · 24-48 hours

Need Clinical-Grade Annotation with Full Provenance?

Tell us about your medical AI dataset and we'll scope a board-certified annotation project with FDA-aligned documentation.

No commitment. NDA available on request. We respond within 24 hours, often the same day for Gulf-region inquiries.

Neel Bennett

AI Annotation Specialist at AI Taggers

Neel has over 8 years of experience in AI training data and machine learning operations. He specializes in helping enterprises build high-quality datasets for computer vision and NLP applications across healthcare, automotive, and retail industries.

Connect on LinkedIn