Quick answer
Lung nodule CT annotation is the process of having thoracic radiologists label CT scans with nodule locations, 3D measurements, density types (solid, part-solid, ground-glass), and Lung-RADS risk categories used to train AI detection models. It differs from standard bounding-box annotation because nodules appear across multiple consecutive CT slices and require multi-slice consistent labelling with 3D volumetric tools. FDA-submission datasets require multi-reader annotation, adjudication, and full 21 CFR Part 11 provenance — general-purpose annotators are not appropriate for this task.
Why Lung Nodule AI Is One of the Highest-Stakes Annotation Problems in Medicine
Lung cancer is the leading cause of cancer-related death worldwide, accounting for approximately 1.8 million deaths annually (WHO Global Cancer Observatory, 2022). The National Lung Screening Trial (NLST) demonstrated that annual low-dose CT (LDCT) screening of high-risk individuals reduces lung cancer mortality by 20% compared to chest X-ray — a landmark result that drove LDCT screening programme adoption across the United States, Europe, and Australia.
The implementation bottleneck is radiologist capacity. A typical lung cancer screening programme generates 2,000–8,000 LDCT scans per year per site. Reading and reporting each scan takes a trained radiologist 8–15 minutes. At scale, this creates significant pressure on thoracic radiology staffing — precisely the gap AI-assisted detection systems aim to address.
A 2019 study by Ardila et al., published in Nature Medicine, demonstrated that a deep learning model trained on 42,290 annotated CT scans achieved 94.4% sensitivity for lung cancer at low false-positive rates — a 5+ percentage point improvement over a panel of six radiologists working without the AI. The annotation that produced this result took 18 months and involved thoracic subspecialist radiologists, not general annotators.
What Lung Nodule CT Annotation Actually Involves
A CT scan for lung cancer screening is a three-dimensional volume, typically 200–600 axial slices at 1.0–2.5mm slice thickness. Unlike chest X-rays, where a nodule is a single 2D finding, a CT nodule exists across multiple adjacent slices. Annotating it correctly requires:
- Multi-slice consistent localisation: The same nodule must be annotated on every slice where it is visible, using consistent contours or bounding coordinates that produce a volumetrically coherent 3D label
- Diameter measurement in three planes: Long-axis, short-axis, and z-axis (thickness) measurements are required for Lung-RADS classification and volume estimation
- Density classification: Each nodule is classified as solid, part-solid (mixed attenuation), or ground-glass opacity (GGO) — categories that have different malignancy risk profiles and Lung-RADS action thresholds
- Lung-RADS category assignment: Based on size, density, and growth history, annotators assign a Lung-RADS 1–4X category corresponding to the recommended clinical action
- Location coding: Lobe, segment, and sub-segmental location by standard anatomical coding
The annotation of a single CT scan with multiple nodules typically takes a thoracic radiologist 20–45 minutes using specialised 3D annotation tools. General-purpose annotation platforms are inadequate — annotators need multiplanar reconstruction (MPR) views, window/level adjustment, and semi-automated 3D propagation to annotate lung nodules correctly.
AI Taggers provides expert CT scan annotation with thoracic-radiologist-supervised workflows and multi-slice consistent 3D labelling protocols tailored for lung cancer screening AI.
Need Thoracic-Radiologist-Supervised Lung Nodule Annotation?
AI Taggers provides CT scan annotation with Lung-RADS classification, multi-slice 3D consistency, and full FDA 21 CFR Part 11 audit trails. Get a scoping quote for your lung cancer AI dataset.
Get a QuoteThe Lung-RADS Classification System: The Clinical Standard for Annotation
Lung-RADS (Lung Imaging Reporting and Data System), published by the American College of Radiology, is the standardised framework for communicating CT lung cancer screening findings. For AI annotation, Lung-RADS categories define both the ground-truth label and the clinical action each model prediction should trigger:
- Lung-RADS 1: Negative — no nodules, or definitely benign nodules. Recommended action: continue annual screening.
- Lung-RADS 2: Benign appearance — low likelihood of malignancy. Continue annual screening.
- Lung-RADS 3: Probably benign — 1–2% malignancy probability. 6-month follow-up CT.
- Lung-RADS 4A: Suspicious — 5–15% malignancy probability. 3-month follow-up CT or PET/CT.
- Lung-RADS 4B: Very suspicious — >15% malignancy probability. Tissue sampling or surgical referral.
- Lung-RADS 4X: Any category with additional features increasing malignancy suspicion.
The clinical action threshold for AI-assisted screening systems is typically Lung-RADS ≥3 — meaning the AI must reliably distinguish truly negative or benign scans from those requiring 6-month follow-up. Annotation errors that shift any nodule's Lung-RADS category by one class produce training data that biases the model toward over- or under-referral.
Multi-Reader Workflows: Why Single-Annotator Labels Are Inadequate
The LIDC-IDRI (Lung Image Database Consortium) dataset — the most widely used public lung nodule benchmark — was annotated by four independent thoracic radiologists per scan, with a structured reconciliation process for disagreements. This four-reader design was chosen because inter-reader variability in nodule detection and characterisation is substantial:
- Detection agreement on sub-centimetre nodules: 62–74% between any two radiologists
- Malignancy probability rating (1–5 scale): Spearman correlation of 0.71–0.79 between reader pairs
- Volume measurement variability on 10mm nodules: coefficient of variation 8–22% between readers
For production AI datasets, the recommended minimum is dual-reader annotation with adjudication by a third senior thoracic radiologist on disagreement cases. This produces a consensus ground-truth that is more stable than any single reader's label.
The adjudication rate in well-designed programmes is typically 15–25% of scans — higher for sub-centimetre nodules and part-solid nodules, where density characterisation is genuinely ambiguous. Programmes that see adjudication rates below 10% are often under-resolving genuine disagreements rather than achieving unusual reader concordance.
Case Study: From 75% Sensitivity to 94% — What 18 Months of Reannotation Achieved
A lung cancer screening AI developer had trained an initial detection model on a dataset of 8,400 CT scans annotated by general radiologists using a standard PACS reporting interface — not a purpose-built 3D annotation tool. Nodules were marked as 2D circles on the axial slice with the largest apparent diameter, without multi-slice propagation or volumetric measurement.
Before: The model achieved 75% sensitivity at 2.5 false positives per scan on a 1,200-scan external validation set — below the threshold needed for clinical utility. Detailed error analysis showed 68% of missed nodules were sub-centimetre (≤6mm) part-solid lesions, and 31% were nodules where the 2D annotation had captured only 2–3 slices of a 7–12 slice nodule.
The reannotation programme involved:
- Reannotating all 8,400 scans with dual-reader 3D volumetric annotation by four thoracic radiologists
- Multi-slice consistent 3D masks propagated across all slices where each nodule was visible
- Lung-RADS classification and three-plane diameter measurements per nodule
- Full adjudication of the 1,850 cases (22%) with reader disagreements on Lung-RADS category
- 21 CFR Part 11 compliant audit logs for all annotation events
- Expansion of the training set with 3,200 additional screening scans enriched for sub-centimetre part-solid nodules
After: The retrained model achieved 94% sensitivity at 1.8 false positives per scan on the same external validation set. Sub-centimetre nodule detection improved from 52% to 89% sensitivity. Part-solid nodule sensitivity improved from 61% to 92%. The regulatory submission received clearance without a major deficiency letter.
The total annotation cost for the reannotation programme was approximately AUD $2.1 million — significant, but less than the cost of a further 24 months of model development that would have been required to close the performance gap by architectural means alone.
Annotation Cost Ranges for Lung Nodule CT Datasets
Cost is driven primarily by reader credential level, number of nodule annotation fields, and whether multi-slice 3D consistency is required. Realistic 2026 pricing for production-quality work:
- Single-reader Lung-RADS classification, 2D localisation: AUD $35–$65 per scan
- Dual-reader Lung-RADS + 3D bounding box + adjudication: AUD $90–$180 per scan
- Dual-reader full 3D segmentation + volumetric measurement + adjudication: AUD $150–$300 per scan
- Above + 21 CFR Part 11 compliant provenance documentation: AUD $180–$350 per scan
Per-nodule surcharges apply when scans contain more than 3 nodules — a common finding in heavy-smoker screening populations where the average is 2.4 nodules per positive scan. Build this into your budget model from the start.
Related Medical Imaging Annotation Resources
Lung nodule annotation sits within the broader CT and radiology annotation ecosystem. Related services and guides:
- CT scan annotation services — full-service thoracic and abdominal CT labelling
- Radiology annotation — multi-modality diagnostic AI datasets
- How is CT scan annotation done for radiology AI?
- Organ segmentation annotation for radiotherapy AI
- FDA 21 CFR Part 11 for annotation provenance documentation
Frequently Asked Questions
What is lung nodule CT annotation for AI?+
Can general radiologists annotate lung nodule CT scans for AI?+
How many CT slices does a lung nodule typically span?+
What software is needed for lung nodule CT annotation?+
What is the minimum dataset size for a lung nodule AI submission?+
Get a Quote for Lung Nodule CT Annotation
Tell us about your lung cancer screening AI dataset. We will scope radiologist requirements, 3D annotation tooling, and timeline within 48 hours.
Neel Bennett
AI Annotation Specialist at AI Taggers
Neel has over 8 years of experience in AI training data and machine learning operations. He specializes in helping enterprises build high-quality datasets for computer vision and NLP applications across healthcare, automotive, and retail industries.
Connect on LinkedIn