Lung cancer screening: fine-tuned MedGemma against twelve Lung-RADS radiologists, more consistent but less sensitive
A team at the French company Median Technologies (Valbonne) compared MedGemma, Google's medical vision-language model, with twelve radiologists assigning Lung-RADS v2022 categories to the CT scans of 481 National Lung Screening Trial patients, 132 of them with confirmed cancer. Used as is, the model reaches an AUC of 0.70; fine-tuned on 7,585 other NLST patients it rises to 0.83, against a mean of 0.90 for the radiologists (0.80 to 0.94 depending on the reader), but then detects only 63% of cancers versus 85%. The paper highlights the model's consistency against human variability; that is a real question, but a consistent model can also be consistently wrong, and everything here stays inside a single twenty-year-old dataset.
The context
Lung cancer screening with low-dose CT reduces mortality in heavy smokers: that is the lesson of the NLST (United States, 2011) and NELSON (Netherlands, 2020) trials. To turn each scan into a decision, the American College of Radiology created Lung-RADS, a scale that assigns each scan a category from 1 (nothing suspicious) to 4X (highly suspicious), each tied to a course of action: follow-up at one year, six months, three months, or biopsy. The scale cuts false positives, but not disagreement between readers: two radiologists regularly put the same nodule in two different categories, and therefore on two different care pathways.
Specialised AI models — Google's 3D network published in 2019, MIT's Sybil in 2023 — have performed well on NLST. What is new is the arrival of foundation models: general-purpose models pre-trained on huge corpora and meant to adapt to many tasks. MedGemma, released by Google in 2025, is a vision-language model (it reads images and produces text) that accepts CT volumes as sequences of slices. The paper's question is simple: can such a model do Lung-RADS classification, what does fine-tuning add, and where does it sit relative to the spread of radiologists?
The method
The data. 481 patients from an NLST subset previously annotated by the same team and published on The Cancer Imaging Archive. 132 cancers (27.4%), confirmed by biopsy within a year of the scan, and 349 non-cancers, confirmed by at least three years of cancer-free follow-up. The cohort is therefore enriched: in the NLST population, the proportion of cancers is about 4%, according to the authors. One scan per patient, thin slices (under 5 mm), from several manufacturers.
The readers. Twelve board-certified radiologists, junior (under three years of practice) and senior, blinded to ground truth and clinical data. Each patient was read by three different radiologists, for 1,443 readings; each radiologist saw about 120 cases. An important point: no prior imaging was available, even though Lung-RADS relies heavily on how a nodule changes over time.
The native model. MedGemma 1.5 in its 4-billion-parameter version, used without retraining (zero-shot). The scan is cut into blocks of 25 consecutive slices with a stride of 5. For each block, a coarse-to-fine prompt first asks for an overall analysis of the lungs, then a description of nodules, then a Lung-RADS category, read from the probabilities the model assigns to each category. The patient score is the maximum across all blocks, a choice that favours sensitivity.
The fine-tuned model. The same model, adapted with QLoRA: instead of retraining all 4 billion parameters, the model is frozen in 4-bit compressed form and only small additional matrices are learned. Training data: 7,585 other NLST patients (393 cancers) with 21,533 hand-annotated nodules. Only blocks containing a malignant nodule are labelled positive, and a focal loss — a loss function that gives more weight to hard examples — offsets the imbalance.
The evaluation. Radiologists and models produce a score on the same ordinal Lung-RADS scale. The authors measure AUC (the probability that a cancer gets a higher score than a non-cancer; 0.5 = chance, 1 = perfect), then sensitivity and specificity at the point that maximises the Youden index (sensitivity + specificity − 1). Confidence intervals come from 5,000 bootstrap resamples. Agreement between radiologists is measured with Fleiss' kappa and the intraclass correlation coefficient (ICC).
The results
Radiologists disagree with each other. Fleiss' kappa of 0.33, "fair" agreement; ICC of 0.68, "moderate" reliability. Seniors agree better (kappa 0.38) than juniors (0.24). On average two readers differ by slightly less than one category.
Discrimination. Radiologists: AUC 0.90 [0.88-0.92], from 0.80 for the weakest to 0.94 for the best. Native MedGemma: 0.70 [0.64-0.75]. Fine-tuned MedGemma: 0.83 [0.78-0.87]. On the roughly 120 cases of the weakest radiologist, the fine-tuned model scores 0.84 versus 0.80, with very wide intervals (0.76-0.92 and 0.67-0.92): the difference is not demonstrated.
The operating point. At the Youden threshold, radiologists have 85% sensitivity and 90% specificity. The fine-tuned model: 63% and 91%. The native model: 75% and 58%. Fine-tuning mainly improves specificity.
Clinical translation. Apply these figures to 1,000 screened people at the roughly 4% prevalence cited by the authors, i.e. 40 cancers. Radiologists would catch about 34 and miss 6, with some 96 false positives. The fine-tuned model would catch 25 and miss 15, with about 86 false positives. The native model would catch 30 but trigger more than 400 false positives. This calculation is optimistic: the thresholds were chosen on the test data itself.
What is good
A real human comparator, with its spread. Twelve radiologists, three independent readings per patient, 1,443 readings in total, blinded, using the clinical scale actually in use. Many papers compare a model with "a radiologist" or an average; here you see the range of human performance and can place the model within it, which is more honest.
Solid ground truth and a reported failure. Cancers are biopsy-confirmed, non-cancers backed by three years of follow-up. And the authors clearly report that MedGemma used as is (0.70) does not reach a clinically useful level: valuable information against the narrative of general-purpose models that work everywhere without adaptation.
Acknowledged limits and open code. The paper itself names the enriched cohort, the lack of external validation, the risk of overfitting to NLST and the absence of ablation studies. The code and protocol are on GitHub, and the annotated subset is deposited on The Cancer Imaging Archive.
What is less good
No external validation, and leakage not ruled out. The model is fine-tuned on NLST and tested on NLST: same era (early 2000s), same centres, same scanners. The paper does not state explicitly that the 481 test patients are disjoint from the 7,585-patient training set, and the authors themselves write that they cannot exclude overfitting. This is the classic setting for data leakage and population bias: an AUC of 0.83 inside one database says nothing about performance on today's scanner fleet in Europe or Asia.
A handicapped comparator and flattering metrics. Radiologists read without prior exams, although nodule growth is central to Lung-RADS: a biased comparator, even if it affects everyone. Above all, the Youden thresholds are optimised on the test data, and the prevalence enriched to 27% makes accuracy misleading. The favourable comparison with the "weakest radiologist" rests on about 120 cases and widely overlapping intervals: a misleading metric if read as equivalence. And 63% sensitivity remains far from the human 85%.
Consistency is not reliability, and the authors mark their own homework. The model always gives the same answer to the same input because its temperature is set to zero: that is mechanical repeatability, not robustness. No test was run with variations in prompt, reconstruction or scanner, as the paper acknowledges. A model that consistently misses one cancer in three is not safer than a variable reader. Finally, all authors are affiliated with Median Technologies, which develops the eyonis screening-support software; no funding or conflict-of-interest statement is detailed in the preprint.
What it changes
For the research community, the paper provides a useful benchmark: a 4-billion-parameter vision-language foundation model, lightly fine-tuned, stays clearly below radiologists on a structured screening task, and the gap shows mainly in sensitivity. Measuring human variability alongside performance is a good idea, but it calls for the symmetric experiment: measure the model's variability under realistic perturbations, and compare it with existing specialised models (Sybil, 3D networks), absent from the analysis.
For clinicians, nothing changes today. A tool that would miss 15 cancers out of 40 where radiologists miss 6 cannot serve as an autonomous reader, nor even as a safety net. The claim of support "in settings with limited expertise" is not demonstrated: it would require a study in which radiologists read with and without the tool.
For patients and the public, the message is cautious: large general-purpose medical models are not ready to interpret a lung screening CT on their own. Disagreement between radiologists is real, and a good reason to look for assistive tools; but a consistent, less sensitive tool is not yet an answer.
Further reading
The paper: Performance vs Consistency: Evaluating a Foundation Model in Lung-RADS Screening, Benjamin Renoust, Pierre Baudot, Tiffany Foriel, Yousra Haddou, Charles Voyton, Pierre-Henri Siot, Ezequiel Geremia, Danny Francis, Jean-Christophe Brisset, Valérie Bourdès, Sylvain Bodard, Benoit Huet (Median Technologies, Valbonne; secondary affiliations: University of Osaka, AP-HP Necker, Memorial Sloan Kettering, Massachusetts General Hospital, Sorbonne Université), arXiv:2609.22281 (cs.CV), submitted 13 September 2026, MICCAI CAPTION 2026 workshop, DOI 10.48550/arXiv.2609.22281. Preprint, not peer reviewed.
The code: EYONIS-AIDS-DS/medgemma-miccai. The base model: MedGemma (Sellergren et al., 2025, arXiv:2507.05201). The scale: ACR Lung-RADS v2022. On low-dose CT lung screening, see also our decryption on synthetic CT degradation and nodule radiomics.