Percival, a vision-language model for CT: what contrastive learning on 436,787 Penn Medicine BioBank scans changes

A University of Pennsylvania team trained Percival, a model that learns to match 3D CT scans with their radiology reports, on 436,787 scans from the Penn Medicine BioBank. Its learned representations predict diagnoses and prognosis better than two comparison models trained without text, but the gain is uneven: Percival is the best of the four approaches on only 279 of the 702 chest CT diagnoses tested. Read it as a promising research model, validated almost entirely within a single health system.

The context

A CT scan (computed tomography) is a three-dimensional volume of millions of points. The AI models that analyze it are usually trained for one narrow task: detecting a lung nodule, segmenting the liver, measuring coronary calcium. Each task requires expert annotations, which are expensive to produce.

A recent alternative is to build a foundation model: a large model trained once, without a specific task, whose internal representations then serve as a starting point for many applications. Two families existed for CT. "Vision-only" models, such as CT-FM, learn from images alone. Multi-organ segmentation models, such as TotalSegmentator, are trained to outline dozens of anatomical structures. Percival explores a third route: using the report the radiologist has already written for each exam as free supervision. This is the CLIP idea, transposed to 3D medical imaging.

The method

Percival is a two-branch vision-language model (VLM). The image branch is a vision transformer (ViT), an architecture that cuts the image into small blocks and learns how they relate to one another. Here the blocks are 16 x 16 x 8 voxel cubes (a voxel is the pixel of a 3D volume), on scans normalized to 352 x 352 pixels in-plane and 128 slices. Three sizes exist: Tiny (6 million parameters), Small (22 million) and Base (86 million). The text branch is CXR-BERT, a language model pretrained on radiology text.

Training is contrastive: for each batch of exams, the model must pull each scan and its own report closer together in a shared mathematical space, and push mismatched pairs apart. It was never told "look for pancreatic cancer"; it only learns what distinguishes one exam from another, relying on what radiologists wrote.

The data come from the Penn Medicine BioBank (PMBB), which covers several hospitals of that health system: 436,787 volumes from 51,966 participants for training, 123,603 volumes from 20,599 participants for validation, for 72,565 participants in total. The split appears to be at the participant level, which limits data leakage between the two sets. The exams span several anatomical regions, contrast phases and scanner types.

The model is then frozen and evaluated by linear probing: only a very simple classifier is trained on top of its representations. If that classifier works well, the useful information was already in the representations. The "labels" are phecodes, groupings of diagnostic codes from the electronic health record: 702 for chest CT and 624 for abdomen-pelvis CT.

The results

ML metrics. On chest CT, the mean AUC (the probability that the model ranks a positive case above a randomly chosen negative one) is 0.84 for circulatory diseases, 0.83 for respiratory and 0.84 for neoplastic, with wide spreads between diagnoses (0.65 to 0.98 for circulatory). A few examples: coronary atherosclerosis 0.85 (versus 0.79 for TotalSegmentator), pancreatic cancer 0.91 (versus 0.84), heart failure with preserved ejection fraction 0.91.

The per-diagnosis count is more informative than the averages. Across the 702 chest phecodes, Percival beats a simple demographic model (age, sex) on 555, beats CT-FM on 512 and beats TotalSegmentator on 315, about 45%. It is the best of the four approaches on 279 diagnoses. On abdomen-pelvis, the figures are close: 511 of 624, 442 and 272, and 230 diagnoses where it ranks first.

For prognosis (Cox models, which estimate the risk of an event over time), the mean concordance index (C-index; 0.5 = chance, 1 = perfect ranking of patients by time to event) is 0.68 on chest CT versus 0.65 for TotalSegmentator, 0.61 for CT-FM and 0.61 for the demographic model. For abdomen-pelvis: 0.70, 0.66, 0.63 and 0.62.

Other analyses: 1,022 associations reaching Bonferroni significance between representations and diagnoses (PheWAS, an analysis across all diagnoses), and 113 associations with 39 laboratory measurements replicated in a second cohort. By contrast, direct prediction of laboratory values remains weak: mean R² of 0.04 across 64 measurements (R² is the share of variance explained), with maxima for albumin (0.40) and hemoglobin (0.30). On external validation, the public CT-RATE dataset (1,564 chest CT scans, 18 labels), the authors report performance "comparable" to a model trained in that domain.

Clinical translation. An AUC of 0.84 says nothing by itself about the number of false positives: that depends on how common the diagnosis is and on the chosen threshold. A purely illustrative example: for a disease present in 1% of patients, a threshold giving 90% sensitivity and 90% specificity would flag, out of 1,000 patients, 9 of the 10 patients with the disease, miss 1, and raise 99 false alarms. The results reported here do not provide this kind of figure, which is normal for a representation model but limits clinical interpretation.

A note on sources: these figures come from the preprint version (medRxiv, version 5). The version published in npj Digital Medicine on 29 September 2026 could not be read in full text for this analysis; differences are possible.

What is good

A rare scale in 3D imaging. More than 436,000 scans paired with their reports, across several hospitals, anatomical regions and protocols, where many CT models are trained on a few thousand exams of a single type.

A broad evaluation that is candid about win counts. The authors test 702 and 624 diagnoses rather than picking three favorable diseases, and the paper lets readers see that the model does not win everywhere. The PheWAS results are corrected for multiple comparisons, and replicating the laboratory associations in a second cohort (113 of 153) is a useful check.

A concrete opening. The code, the weights of the three sizes, the training pipeline and the downstream classifiers are published on GitHub and Hugging Face; a version 2.0 came out in May 2026. The repository also includes evaluation on CT-RATE, which eases reproduction on public data.

What is less good

Population bias and thin external validation. The data come from a single US health system, and the only external validation reported is thoracic (1,564 scans), whereas the results also cover abdomen-pelvis CT and more than 1,300 diagnoses. The authors themselves note the absence of public cohorts with laboratory data and longitudinal diagnoses. Detailed demographics and exclusion criteria do not appear in the excerpts we could consult.

Shortcut learning and leakage through language. The model is trained to align with radiologists' text. The authors acknowledge that what comes from the image cannot be separated from what comes from the vocabulary used in the reports. Reports mostly describe normal exams, and a single report sometimes summarizes several volumes from the same visit. Two further risks apply: index-event bias (exams are ordered because a problem is suspected, so diagnoses close to the exam are partly already known), and labels derived from record codes rather than adjudicated diagnoses.

Comparators and metrics that limit the scope. The comparators are other paradigms (vision-only, segmentation), not other CT vision-language models; and the mean prognostic gain (C-index 0.68 versus 0.65 for TotalSegmentator) is modest. Follow-up is not specified (censoring at the last record contact), no hazard ratios appear, and few confidence intervals are given. Finally, the code license is CC BY-NC 4.0, non-commercial, which rules out production use without a separate agreement. On funding: NIH and Penn Medicine foundations; one author declares consulting activities with pharmaceutical companies.

What this changes

For the research community, Percival provides a public, documented starting point for text-aligned CT models, with a large-scale evaluation protocol and downstream classifiers already released. Next studies will need to test transferability to other hospitals and pit Percival against other vision-language models.

For clinicians, nothing changes today: this is a research tool, with no prospective validation or regulatory status. In the longer term, extracting a risk signal from a scan done for another reason (so-called "opportunistic" reading) is plausible, but a benefit for the patient will have to be shown, not just an AUC.

For patients and the public, the message is cautious: a CT scan holds more information than a report retains, and models are starting to use it. Whether that information improves care remains to be demonstrated, hospital by hospital.

Further reading

The paper (published version): Generalizable CT vision-language modeling for population health and disease risk, Cameron A. Beeche, Joonghyun Kim, Walter R. Witschey et al. (University of Pennsylvania), npj Digital Medicine, 29 September 2026, DOI 10.1038/s41746-026-03257-2. Preprint: medRxiv, version 5. Code and weights: github.com/cams2b/percival (CC BY-NC 4.0).

On contrastive alignment between two modalities, see our decryption on ECG and cardiac MRI in Chagas disease. On the reliability of imaging encoders, see our decryption of CRS-Bench.