One model for the whole PSMA PET/CT read: report, Q&A and segmentation — with evaluation labels written by Gemini
Fourteen authors, led by the University of Florida, assemble a single vision-language model that writes the PSMA PET/CT report, answers questions about it and segments lesions, trained on 5,747 in-house studies and on the PSMA subset of AutoPET. It beats its comparators everywhere: BLEU 36.2 against 27.4 for PET2REP, 83.6% correct on multiple-choice questions against 70.9%, lesion Dice 0.588 against 0.550 for nnUNet. But the 10,274 test questions were written and then verified by Gemini-2.5-Pro with no human ever reading them, no nuclear medicine physician read a single model output, and the voxel-level grounding claimed in the abstract is — by the authors' own admission — neither trained nor evaluated.
The context
PSMA PET/CT has become, in a handful of years, the pivotal examination in prostate cancer. The acronym stands for prostate-specific membrane antigen, a protein heavily overexpressed on the surface of cancerous prostate cells. The patient is injected with a molecule that binds to it, labelled with a positron-emitting isotope — here 18F-DCFPyL — and the PET camera shows where the tracer accumulates, while the CT provides the anatomical reference. The result is a whole-body study far more sensitive than conventional imaging for staging and for finding recurrence after treatment.
That sensitivity has a reading cost. A single study may contain dozens of foci scattered from the pelvis to the axial skeleton, and the physician must separate lesions from physiological uptake — the tracer normally fixes in the kidneys, bladder, liver, spleen and salivary glands, sometimes more intensely than a metastasis. Reading takes time, and inter-reader reproducibility is imperfect.
AI applied to this modality has so far advanced task by task: segmentation models that outline lesions — the public AutoPET challenge structured that work — report-generation models such as PET2REP, and visual question answering models (VQA: ask a natural-language question about an image, get a text answer). Each with its own data, pipeline and metrics. The paper's argument is architectural: a human reader does not separate these gestures, they describe what they see and point at what they describe. A single model could therefore ground its text in the voxels it segmented. That is the promise; what follows is whether it holds.
The method
The language data. A retrospective University of Florida cohort: 5,747 whole-body PSMA PET/CT studies acquired between 2022 and 2025, all with 18F-DCFPyL, each paired with its clinical report converted into structured JSON (date, findings, impression). IRB approval with waiver of consent. The split is 4,023 / 574 / 1,150 studies for training, validation and test — and the paper states it is done "at the case level". No demographics are reported: no age, no PSA, no stage, no indication. No inclusion criteria, no count of distinct patients, no scanner details.
The segmentation data. The PSMA subset of the AutoPET challenge: 597 studies in 378 male patients from LMU University Hospital Munich, acquired between 2014 and 2022, with lesion masks manually annotated in 3D on the PET images. Patient-level split this time: 419 / 59 / 119. Both cohorts undergo the same spatial normalisation — resampling to 2.73 × 2.73 × 2.8 mm, a fixed 256 × 256 × 400 voxel matrix, field of view stopping at the upper thighs, PET converted to SUV (standardized uptake value, the unit that normalises uptake by injected dose and patient weight).
The architecture. The assembly follows the LLaVA blueprint, now standard for vision-language models. A 3D vision encoder (a two-channel vision transformer, PET and CT) is first pretrained contrastively in CLIP fashion: the model learns to pull a volume close to its own report and away from every other. An MLP-Mixer projection module then compresses the visual sequence — far too long in 3D for a language model — into 512 tokens of dimension 2,048. Those tokens are fed into a LoRA-tuned language model (low-rank adaptation: the original weights are frozen and only small additional matrices are trained, here of rank 128). Finally an nnUNet-style 3D segmentation branch receives, through a special [SEG] token and a LISA-style cross-attention mechanism, the language model's internal state — the bridge meant to link text to voxels. Auxiliary anatomical supervision, derived from TotalSegmentator over ten organ groups (kidneys, bladder, liver, spleen, brain, lungs, heart, prostate, stomach, salivary glands), is added to teach the model not to confuse physiological uptake with disease.
One point deserves stating immediately: the language model used is never named. Not its maker, not its size, not its version. The architecture figure reads "LLM Backbone" and the text says no more.
How the questions were built. This is the decisive methodological point. The VQA set was not written by physicians: it was generated by the Gemini-2.5-Pro API from the de-identified reports, following a recipe borrowed from CT-CHAT. Nine items per study — three description, three multiple-choice, three free-response — on imposed themes: disease presence, distribution, local recurrence in the prostate bed, pelvic and extra-pelvic nodes, metastases. A second pass asks the same model to verify that each answer is supported by the report, identifying the supporting text; non-compliant items are rejected or regenerated. What remains is 35,838 training questions, 5,166 validation and 10,274 test. No human review is reported at any stage.
Training runs in four stages: contrastive pretraining of the vision encoder; alignment of the projection layer alone, vision and language frozen; fine-tuning of the vision-language branch on report and VQA; then final multitask tuning where text and segmentation are optimised jointly, with an autoregressive loss on one side and cross-entropy plus Dice on the other. The hardware is detailed (eight NVIDIA B200 GPUs, PyTorch 2.7) but no hyperparameter is: no learning rate, no batch size, no epoch count, no task weighting. Only the optimiser is given, AdamW.
The results
Image-text retrieval. The pretrained encoder is first evaluated on its own, on its ability to find the right report from a volume. Across the 1,150 test studies, the correct report appears in the top 100 results in 35.6% of cases [95% CI: 32.6–38.3], against 14.2% for a CT-CLIP trained on CT alone and 7.5% for random draw. The gap is clear and cleanly measured: PET/CT-specific pretraining delivers something real.
Report generation. The proposed model reaches ROUGE 23.4, BLEU 36.2, METEOR 19.6, CIDEr 3.13 and BERTScore 86.1. The comparators are PET2REP (ROUGE 16.7, BLEU 27.4, METEOR 17.0, CIDEr 0.53, BERTScore 84.2) and MedVL-SAM2, a multimodal baseline that sees only the CT (ROUGE 12.5, BLEU 9.1, CIDEr 0.015). All bootstrap confidence intervals are disjoint. But one must know what these metrics measure: BLEU, ROUGE, METEOR and CIDEr count overlaps of word sequences between generated and reference text. They say nothing about diagnostic accuracy. A report writing "no bone lesion" where the reference writes "bone lesion" loses two words in a hundred of BLEU, and a patient in the clinic. No clinical factual-accuracy metric — RadGraph F1, GREEN, RaTEScore, all available today for this kind of evaluation — is reported. BERTScore is not one: it measures semantic proximity, not truth.
VQA. On multiple-choice questions the model reaches 83.60% correct [82.39–84.87] against 70.89% [69.29–72.39] for MedVL-SAM2. On free responses CIDEr goes from 17.5 to 91.6. Translation: one multiple-choice question in six is still wrong, on a question set whose difficulty was set by another language model.
Segmentation. Dice — the overlap between predicted and annotated volume — reaches 0.5875 [0.5426–0.6314] for the language-guided variant, 0.5858 for the semantic variant, against 0.5498 [0.4962–0.6012] for nnUNet and 0.4653 for SegAnyPET. The authors add three lesion-level F1 scores with three matching criteria: any overlap (0.736 against 0.653 for nnUNet), coverage of the maximum-SUV voxel (0.727 against 0.637), and lesion Dice above 0.5 (0.593 against 0.492).
In clinical terms. A lesion-level F1 of 0.736 on the most permissive criterion means that roughly a quarter of the lesions counted are either missed or invented. A Dice of 0.59 means the outlined volume covers the true volume only 59% of the way: usable for plain detection, not usable for lutetium-177 radioligand dosimetry, where tumour volume drives the dose. And on Dice specifically, the proposed model's confidence interval overlaps nnUNet's substantially — the 0.038-point gain is not established. No paired statistical test is reported anywhere in the paper: the statistics section is limited to bootstrap intervals, with the number of resamples unstated.
What is good
The contrastive PET/CT pretraining is justified by a measurement, not an intuition. The authors could have simply asserted that a dedicated encoder beats a recycled CT one. They demonstrate it on an intermediate task, image-text retrieval, with a relevant comparator (CT-CLIP) and an explicit floor (chance). CT-CLIP turns out to be barely above chance on text-to-image retrieval, which is in itself a useful result for the field: reusing a CT encoder on nuclear medicine does not work.
The segmentation metrics are honestly chosen. Global Dice is a poor measure in PSMA PET, because it is dominated by the few large lesions and ignores the small foci that actually flip staging. Reporting three lesion-level F1 scores, one of them anchored on the maximum-SUV voxel, is the right reflex. And all three criteria are given together, including the harshest, on which the gap to nnUNet is widest.
The anatomical supervision targets a real failure mode. Adding ten TotalSegmentator organ groups is not decoration: renal, bladder and salivary physiological uptake is the leading source of false positives in PSMA reading. Giving the model an anatomical map to learn to discount them answers a problem identified by clinicians, not an architectural need. The authors also list their main limitations without padding, and declare their funding (NIH R01EB034692 and R01AG078250) and the absence of competing interests.
What is less good
The VQA evaluation is circular. This is the central limitation. Gemini-2.5-Pro writes the questions from the reports, then Gemini-2.5-Pro verifies that its own answers are supported by those same reports. The 10,274 test questions constitute the ground truth, and not one was reviewed by a physician. What the 83.6% score measures is therefore the agreement between a trained model and another model judging its own output. The decisive control is also missing: no image-blind baseline. A multiple-choice question derived from a report is often guessable without the image, through internal consistency or answer distribution — a textbook case of shortcut learning, where the model learns the spurious correlation rather than the task. Until a language model alone, deprived of any image, has been run on the same set, we do not know what share of the 83.6% comes from reading the PET.
There is no external validation, and the split leaves the door open to data leakage. All three language tasks are trained and tested on the same cohort, one centre, one tracer, one period. No other hospital, no other scanner, no test on 68Ga-PSMA-11, which is widely used elsewhere in the world. More worrying: the paper explicitly states a split at the case level, not the patient level. Yet a man followed for prostate cancer typically undergoes several PSMA PET/CT studies — initial staging, post-treatment control, recurrence work-up. If the same patient lands on both sides of the fence, that is data leakage in the strict sense, and the reported performance is optimistic by an unknown amount. The paper does not give the number of distinct patients behind the 5,747 studies, which prevents even estimating the size of the risk. As for AutoPET, it is used in training and test: that is not external validation, it is a second source internal to the pipeline. No cross test — train in Florida, evaluate in Munich — is attempted.
No human read a single model output, and the grounding promise is not tested. No reader study, no error analysis, no hallucination count: the authors explicitly defer this to future work. For a system whose main output is a report meant for a physician, that is the evaluation that is missing, and no factual metric replaces it. Above all, the argument that justifies the unified architecture — the voxel-level grounding announced in the abstract — does not exist in the strong sense, and the authors acknowledge it: the report and segmentation datasets are not paired at the case level, so nothing forces a segmented lesion to be mentioned in the text, nor a lesion described in the text to correspond to a mask. The paper's thesis is asserted in the abstract and invalidated in the discussion.
Add reproducibility gaps that are not details: the language model is unnamed, no training hyperparameter is given, and there is no ablation study on an architecture made of five distinct blocks — it is impossible to know what the MLP-Mixer, the TotalSegmentator supervision, the cross-attention or the LoRA rank contribute. Code and weights are not released, despite a section titled "Code availability" that contains only library version numbers. The in-house data is available only on request to the corresponding author.
What it changes
For the research community, the paper has mapping value. It shows that a standard LLaVA assembly transposed to 3D PET/CT stands up and beats single-task models on their own metrics, and it identifies the bottleneck precisely: what is missing is not architectures but datasets where the report and the lesion masks describe the same studies. As long as those two sources stay disjoint, text-to-voxel grounding will remain an illustrative promise. The second lesson is negative and just as useful: LLM-generated VQA benchmarks, now common practice because they are cheap, produce scores nobody knows how to interpret without an image-blind baseline and a clinician-reviewed sample alongside.
For clinicians, nothing changes today. No reader study, no external validation, no code, no weights, no regulatory clearance, a single tracer. What deserves attention on a three-to-five-year horizon is the direction: a reading assistant you can question and that answers by showing the voxels it relies on would be, unlike an opaque detection score, verifiable at the bedside. Provided the grounding is actually trained — and that someone has measured how often it is wrong.
For patients and the general public, the lesson is one of vocabulary. A headline announcing that "AI now reads prostate PET scans better than existing models" would be literally accurate and practically misleading: the model beats other models on word-overlap metrics and on questions written by another AI, in a single American hospital. No physician has yet checked that a single one of its reports was correct.
Further reading
The preprint: A Unified Vision-Language Model for PSMA PET/CT Report Generation, Visual Question Answering, and Lesion Segmentation, arXiv:2609.15603, submitted 14 September 2026, CC BY 4.0, 24 pages. Authors: Yang Xing, Jiong Wu, Savas Ozdemir, Yang Zhou, Boxiao Yu, Ying Zhang, Zheren Zhu, Chenyu You, Wei Shao, Yang Lu, Kang Wang, Tinsu Pan, Yang Yang and Kuang Gong (corresponding), for the University of Florida at Gainesville and Jacksonville, the University of California San Francisco, Stony Brook University and the UT MD Anderson Cancer Center. Funding: NIH R01EB034692 and R01AG078250; no competing interests declared. Neither code nor weights are released; in-house data is available on reasonable request to the corresponding author. The segmentation set is public: AutoPET.
On the same questions, see our decryptages on UNETR segmentation of prostate cancer in PSMA PET/CT and survival prediction, on clinical evaluation of generated pathology reports, and on the circularity of benchmarks whose labels are written by a language model.