Segmenting the pancreas on CT and MRI alike: what adversarial modality alignment shows, and what it never measures

A team from the radiology department at Northwestern University, with two gastroenterologists from Northwestern and Mayo Clinic Florida, trained a single network to segment the pancreas across 4,604 CT and MRI scans combined, adding one explicit constraint: the model must no longer be able to tell which modality an image came from. Overlap reaches 87.31% on the internal test set and stays between 84.20% and 88.09% across four external datasets, and the representations learned this way allow the pancreas to be split into head, body and tail on CT scans even though no CT subregion annotation was ever seen. The problem lies elsewhere: the paper's two results tables contain only the proposed method, with no numerical comparator, no standard deviation, no statistical test — and the finding presented as most notable rests on seventeen scans annotated by one reader.

The context

The pancreas is the abdominal organ segmentation algorithms miss most often. It is retroperitoneal, meaning pressed against the back of the abdomen behind the stomach; its shape and position vary widely between people; and its contrast against the duodenum, spleen and surrounding vessels is poor. Segmenting, here, means assigning every voxel — the three-dimensional pixel of a volumetric scan — a "pancreas" or "not pancreas" label. It is the building block for everything else: measuring a volume, tracking a cyst, drawing a contour before radiotherapy, extracting texture descriptors.

Two modalities coexist in practice. CT remains the reference in pancreatic oncology. MRI is gaining ground for surveillance of cystic lesions, notably IPMNs, because it involves no radiation and these scans are repeated for years. The two images have nothing in common in terms of intensities: on CT, a voxel value is a physical, reproducible Hounsfield unit; on MRI it depends on the sequence, the machine and the settings, and has no absolute unit.

The usual answer is to build two models, one per modality. That is expensive in annotation, and it learns nothing shared about anatomy. Naively mixing the two datasets generally degrades performance: the network starts exploiting each modality's intensity signature rather than the shape of the organ — a textbook case of shortcut learning, where a model latches onto a correlate of the answer rather than the structure being sought. The domain adaptation literature has long offered a remedy, the domain-adversarial training introduced by Ganin and colleagues, adapted to medical imaging by Kamnitsas on brain lesions and then by SIFA for CT-MRI transfer. The authors also cite PaNSegNet as the current state of the art for the pancreas on MRI. Remember that name: it will not come back.

The method

A shared encoder, a discriminator, and a layer that reverses gradients. The backbone is a standard 3D nnU-Net — the encoder-decoder architecture that remains, in 2026, the hard-to-beat reference in medical segmentation. The encoder compresses the volume into an abstract representation; the decoder expands it back into a mask. Here a single encoder-decoder is shared by both modalities, and the segmentation loss combines a Dice term and a cross-entropy term.

The contribution is attached to the network's bottleneck, where the representation is most compact. A small classifier, the domain discriminator, is added there with one job: guess whether the image is CT or MRI. Between the encoder and this discriminator, the authors insert a gradient reversal layer — a device that passes information through unchanged on the way in, but flips the sign of the gradient on the way back. The consequence: the discriminator learns to tell modalities apart while the encoder is pushed to make it fail. At equilibrium, the representation no longer carries a modality signature. The total loss weights the two objectives by a factor lambda, initialised at 0 and raised gradually to 0.5; the exact schedule of that ramp is described by the authors themselves as "empirically designed".

Training. Two stages: 2,000 epochs of segmentation alone on mixed CT and MRI, then 1,000 epochs with the discriminator. Patches of 80 × 160 × 192 voxels, Adam optimiser at 0.001, z-score normalisation, sliding-window inference, on a server with eight A6000s. Parameter count, discriminator architecture, batch size and augmentation details are not reported.

The subregion transfer. This is the second half of the paper, and the idea is clean. Once the encoder is trained, it is frozen and a new decoder is attached, trained to split the pancreas into head, body and tail — but only on MRI labels, those of the Cyst-X dataset, the only one of the four training sets to provide that split. Then the whole thing is applied to CT scans. If the encoder really is modality-agnostic, a decoder trained on MRI should work on CT. No operational anatomical definition of the three subregions is given — no cutting plane, no vascular landmark such as the superior mesenteric vein.

The data. 4,604 in-distribution scans: 1,461 Cyst-X MRIs, 1,888 MRIs from an undescribed private set, 1,000 AbdomenCT-1K CTs, 255 CTs from a "peri-pancreatic edema" set. Three MRI sequences are represented — T1-weighted, T2-weighted and out-of-phase — with no numerical breakdown. The whole is split 8:1:1 into training, validation and test, with the text never stating whether that split is at the patient or the scan level. To this are added 440 scans declared out of distribution: AMOS in MRI (60) and CT (300), U-Mamba in MRI (50), BTCV in CT (30). No demographics, no inclusion or exclusion criteria, no list of centres despite the "multi-center" claim, and nothing about the private set that nonetheless makes up 41% of the corpus.

The results

Whole pancreas. On the internal test set, overall Dice of 87.31%, HD95 of 5.23 mm, ASSD of 0.94 mm, IoU of 78.42%. By subgroup: T1W at 85.59, T2W at 88.27, out-of-phase at 89.76, CT at 87.53. On the four external sets: AMOS-CT at 84.20, BTCV-CT at 84.58, AMOS-MRI at 86.87, U-Mamba-MRI at 88.09. Dice measures overlap between the predicted mask and the reference mask; HD95 measures, in millimetres, the contour error at the 95th percentile — how far away the most wrong points sit once outliers are set aside. The two do not say the same thing, and the paper offers a fine illustration: the out-of-phase sequence has the best Dice in the whole test set (89.76) and the worst HD95 (8.47 mm, nearly triple T2W's 2.92 mm). Good volumetric overlap with locally very wrong contours — typically a fragment of a neighbouring structure picked up far from the organ.

Subregions. On MRI, an average of 80.53%: head 85.19, body 80.23, tail 76.18. On CT, with no subregion label seen in training, an average of 83.05%: head 82.29, body 83.26, tail 83.61. This is the result the authors present as most notable, and it rests on seventeen scans — 7 from AMOS and 10 from BTCV — hand-annotated for the occasion by a single expert radiologist, with no second reading and no inter-observer agreement measure.

What the tables do not contain. No standard deviation, no confidence interval, no statistical test, anywhere. And above all, no comparison row. No nnU-Net without the discriminator, no lambda at zero, no single-modality model, no PaNSegNet. Tables II and III contain only the proposed method.

Clinical translation. A Dice of 76% on the pancreatic tail means roughly a quarter of the delineated volume disagrees with the reference. For a glandular volume measurement in diabetes research, that is acceptable; for drawing the body-tail boundary that separates the indication for a pancreaticoduodenectomy from that for a distal pancreatectomy, it is nowhere near a surgical margin. The paper claims nothing of the sort: it has no clinical endpoint, no clinician assessment of whether the contours are usable, no measurement of correction time.

What is good

The scale and heterogeneity are real. 4,604 internal scans and 440 external ones, two modalities, three MRI sequences, four named and publicly available external datasets. Much pancreatic segmentation work still makes do with a few hundred CTs from a single centre. Here the "multi-center" claim is undocumented, but the volume and sequence diversity are not in doubt.

The subregion transfer test is honestly designed. The authors could have settled for a qualitative figure showing tidy contours on CT. Instead they had seventeen volumes hand-annotated to produce a number. The sample is tiny, but the methodological intent is right: you do not claim a transfer without going and measuring it in the target modality.

Distance metrics are published alongside Dice, and broken down. HD95, ASSD and IoU appear sequence by sequence and dataset by dataset, not only in aggregate. That is what makes the out-of-phase anomaly above visible at all. Publishing numbers that partly contradict your own narrative is a virtue, even when the text does not comment on them.

What is less good

The central contribution is never measured. This is the dominant weakness, and it is structural. The paper asserts that adversarial alignment produces modality-independent representations and that this is what delivers the performance. The only evidence offered is a t-SNE projection — a two-dimensional visualisation of the latent space — which the authors themselves concede is "qualitative". An ablation exists, but it covers only the downstream fine-tuning, its values live in a figure alone, and its wording is already a concession: the pretrained model "outperforms or is comparable to" a baseline never defined in the text. "Comparable to" is not a result. This is an absent comparator rather than a biased one, a more radical variant of the same flaw. PaNSegNet, named as state of the art a few paragraphs earlier, is work from the same laboratory on largely overlapping data: its absence from the table is hard to explain other than as a decision not to run it.

The risk of data leakage is not ruled out. The 8:1:1 split is announced without a word about patient-level stratification. Yet Cyst-X provides T1W, T2W and out-of-phase for the same patients, and the results table reports performance by sequence. A scan-level draw would put one patient's T1W in training and their T2W in test — same anatomy, same cyst, same glandular morphology. This is the most mundane failure mode in the health-AI literature, the one that collapses published scores as soon as it is properly controlled, and also the one a three-word phrase would dispel. That phrase is not there. A related doubt hangs over the "out of distribution" status of AMOS-CT and BTCV: AbdomenCT-1K, used in training, is by construction an aggregation of public abdominal CT datasets, and no deduplication procedure is described. The same dataset-origin signature mechanism has already been documented in mammography.

The headline finding compares two populations, not two modalities. CT beats MRI on subregions — 83.05 against 80.53 — despite having seen no subregion label. The authors do not comment on the inversion. A simple explanation is available and goes undiscussed: the MRI test comes from Cyst-X, a cohort assembled to study cystic lesions, hence pathological and distorted pancreases; the CT test comes from AMOS and BTCV, abdominal organ segmentation sets that are mostly normal. The gap therefore measures case difficulty as much as model capability: a population bias compounded by a misleading metric. The same criticism applies to the 87.31% headlined in the abstract: it aggregates 3,349 MRIs and 1,255 CTs with no explicit weighting, and so mostly reflects MRI performance. Add to this the total absence of cohort characterisation — no age, no sex, no disease status, nothing on the 1,888 private MRIs — which makes any subgroup analysis impossible, including by proxy.

What it changes

For the research community, the idea is worth picking up, independently of how it is demonstrated here. A modality-agnostic encoder is a reusable object: it turns a stock of annotations available in one modality into usable supervision in the other, and annotating pancreatic subregions on CT is exactly the kind of work nobody wants to do twice. But the question this paper does not answer is the only one that decides adoption: does adversarial alignment add anything an nnU-Net trained on both modalities mixed does not already deliver? One extra row in Table II would settle it. As long as it is missing, the reader cannot separate the effect of the method from the effect of 4,604 scans.

For clinicians, nothing today. Code and weights are announced "upon acceptance", meaning unavailable, and the preprint itself is under a CC BY-NC-ND licence that forbids derivatives and commercial use. There is no prospective validation, no radiologist review of the contours produced, no measurement of how long correcting them would take. The underlying question — is an automatic contour good enough that an operator gains time rather than losing it — is not asked, even though it is the real adoption criterion and we know how to measure it, as other work on operator workload has done.

For patients and the public, a useful distinction: segmenting is not diagnosing. A Dice of 87% does not mean "87% correct diagnoses", it means the machine's outline overlaps a human's by 87%. It is plumbing, essential and invisible, upstream of everything else. The interesting question for a patient under surveillance for a pancreatic cyst is not the model's Dice but whether the volume measured this year is comparable to the one measured last year on a different machine — a reproducibility question this paper makes plausible without measuring it.

Further reading

The preprint: Unified CT and MRI Pancreas Segmentation for Label-Efficient Cross-Modality Subregion Transfer, arXiv:2609.13043, submitted 11 September 2026 under cs.CV, DOI 10.48550/arXiv.2609.13043, under a CC BY-NC-ND 4.0 licence. Authors: Ziliang Hong, Hongyi Pan, Halil Ertugrul Aktas, Andrea Bejar, Elif Keles, Frank H. Miller, Michael B. Wallace, Rajesh N. Keswani, Gorkem Durak and Ulas Bagci, Department of Radiology at Northwestern University, with the Division of Gastroenterology and Hepatology at Northwestern and at Mayo Clinic Florida.

Code and weights: announced as released "upon acceptance", hence unavailable to date; no repository is given. Public datasets used: AbdomenCT-1K, AMOS, BTCV, U-Mamba. Access terms for Cyst-X and the "peri-pancreatic edema" set are not specified in this manuscript, and the private set of 1,888 MRIs is not shared. No mention of ethics committee approval appears in the text.

Declared funding: NIH grants U01-CA268808 and NHLBI R01-HL171376. No conflict-of-interest statement appears in the manuscript, even though two of the authors are practising clinical gastroenterologists.

On the same ground, the works cited as antecedents: Ganin and colleagues' DANN for gradient reversal, Chen and colleagues' SIFA for CT-MRI adaptation, Isensee and colleagues' nnU-Net for the backbone, and the same group's PaNSegNet for the state of the art in pancreatic MRI.