AI-generated thin-slice CT: training on real pairs rather than simulated thick slices changes which method wins
A team from Northeastern University in Shenyang, the US National Cancer Institute and three Chinese hospitals releases MSS-CT, a set of 2,000 chest CT exams in which thick-slice and 1 mm series were reconstructed from the same raw data, together with BasicCSS, a reference model that synthesises the missing thin slices. The finding that matters is not that BasicCSS comes out on top, but that training such models on simulated thick slices — the dominant practice in this literature — degrades every method tested and goes as far as flipping their ranking. Validation, however, stops at image similarity, at the consistency of automated measurements, and at two unblinded radiologists: no diagnostic endpoint was measured.
The context
A CT scanner does not produce images directly: it records raw projections, which are then reconstructed into a stack of slices according to chosen parameters — slice thickness, slice interval, reconstruction kernel. Slice thickness sets the resolution along the head-to-feet axis. A 1 mm series samples the volume finely; a 5 mm series aggregates five times more tissue per voxel and suffers from the partial volume effect: when a single voxel mixes vessel, air and bronchial wall, contrast collapses and small structures blur.
So why keep thick slices at all? Because thin slices are noisier at equal dose, multiply storage and transfer volume fivefold, and many departments — in screening, on call, in budget-constrained health systems — archive only the thick series. The radiologist who wants to measure a nodule or judge a bronchial wall is then left without the series they need.
The conventional answer is interpolation: resampling the 5 mm stack to 1 mm spacing. This standardises geometry but recovers nothing — the information was lost at reconstruction, not at display. Hence the idea of deep-learning CT slice synthesis (CSS): a model learns to infer plausible intermediate slices from thick-slice inputs.
The methodological problem this paper attacks lies elsewhere. Training such a model requires (thick slice, thin slice) pairs from the same patient. Because these are scarce, the literature manufactures pseudo-pairs: take a real thin series and artificially degrade it to imitate a thick one. Except that a real 5 mm slice is not a blurred thin slice: it results from signal aggregation in the reconstruction domain, with the vendor's kernel and characteristics. Training on an imitation and deploying on the real thing is the same gap already documented in synthetic degradation of low-dose CT.
The method
The dataset. MSS-CT gathers 2,000 subjects in two subsets. MSS-CT 3 mm: 500 paired 3 mm / 1 mm exams from the General Hospital of Ningxia Medical University, collected November 2023 to December 2024 on a Philips scanner, median age 64.5 years (interquartile range 52.0-71.0), 45.6% male, clinical status not recorded; split 350 / 50 / 100. MSS-CT 5 mm: 1,500 paired 5 mm / 1 mm exams from the Second Affiliated Hospital of Shenyang Medical College, March 2023 to June 2025, Neusoft scanner, median age 34.0 years (24.0-61.0), 57.7% male, 1,000 healthy subjects and 500 abnormal; split 1,050 / 150 / 300, stratified by health status. The decisive point: in both cases the two series come from the same raw projections, which makes registration unnecessary and guarantees voxel-to-voxel alignment.
External validation. Two independent sets, for the 5 mm to 1 mm task only: RPLHR-CT, a public set of 250 subjects (Philips), and EC-CT, an institutional set of 300 subjects from the Chinese PLA General Hospital (Philips, median age 25.0 years, 200 healthy / 100 abnormal). Models trained on MSS-CT 5 mm were tested there without fine-tuning.
The model. BasicCSS chains a volumetric encoder, a slice expansion module that stretches features along the axial direction to the target resolution, a block-wise decoder that locally refines each interval between two adjacent thick slices, a cross-view refinement module that corrects sagittal and coronal continuity through multi-reference attention, and a reconstruction tail. The basic blocks combine convolutions with Swin-Transformer layers — a transformer whose attention is computed over sliding windows, cheaper than global attention on a 3D volume. Training is simple and fully reported: intensities clipped to [−1024, 2048] Hounsfield units then normalised, L1 loss against the real 1 mm series, AdamW at 10⁻⁴, batch size 1, patches of 8 × 128 × 128 voxels, on NVIDIA RTX A6000 GPUs.
The comparators. Sixteen methods spanning the three families of slice synthesis (reformatted-view super-resolution, volume super-resolution, slice-wise interpolation), plus natural-image super-resolution and video frame interpolation models adapted to the task, plus trilinear interpolation as the conventional reference. Three representatives are discussed in the main text: ACVTT, ResVox and CTHNet. The authors state that model sizes were kept within comparable ranges.
The four evaluation axes. (1) Image fidelity, via PSNR — peak signal-to-noise ratio, in decibels, measuring pixel-wise deviation from the reference image — and SSIM, the structural similarity index computed separately in the axial, coronal and sagittal planes. (2) Effect of training-pair type: real pairs against two pseudo-pair variants (thin-slice sampling, or axial degradation). (3) Downstream consistency: automated segmentation with TotalSegmentator v2.11 of ten thoracic structures (heart, aorta, lung vessels, trachea, T3 and T7 vertebrae, bilateral fifth and seventh ribs), compared by Dice coefficient against the segmentation obtained on the real 1 mm series; and radiomic consistency, measured by concordance correlation coefficient (CCC) on randomly sampled lung regions of interest. (4) A reader study.
The results
Image fidelity. BasicCSS achieves the best PSNR and the best SSIM in all three planes, on both tasks, internally and externally, with all paired comparisons significant (two-sided Wilcoxon tests with Bonferroni correction, p < 0.001). On the 5 mm to 1 mm task, BasicCSS reaches 44.09 ± 1.43 dB against 42.11 to 43.48 dB for the other synthesis methods. The gap to the strongest comparator, CTHNet, is 0.29 dB on the 3 mm task and 0.61 dB on the 5 mm task — roughly a 7% reduction in root-mean-square error. On the 3 mm task, trilinear interpolation tops out at 37.97 ± 0.93 dB with SSIM of 0.967 / 0.964 / 0.963. One caution the paper does not state explicitly: absolute PSNR values are not comparable across subsets, since vendors, protocols and populations differ; only within-set comparisons are meaningful.
The central result: pseudo-pairs mislead. For the six methods tested and on both tasks, real-paired training yields higher PSNR than both pseudo-paired settings, with a larger gap on the 5 mm task. Axial degradation does better than plain slice sampling, but both stay below real pairs. Above all, the ranking flips: under pseudo-pairs ACVTT comes first; under real pairs BasicCSS does. In other words, a method ranking established on simulated data can crown a different winner than the one obtained under clinical conditions — which retrospectively invalidates part of the published comparisons.
Downstream consistency. Across all four datasets and ten structures, segmentations derived from synthesised slices are closer to the reference than those derived from interpolation (p < 0.01 throughout), with a larger gain on the 5 mm task. The structures that suffer most from interpolation are lung vessels, the trachea and bone. On the radiomic side, features extracted from synthesised slices agree better with real 1 mm features than those from the thick slice or from interpolation, with the reproducibility threshold set at CCC ≥ 0.85. The benefit is clear for LoG features (Laplacian of Gaussian filtering, sensitive to density transitions) and weak for wavelet features, which remain the most sensitive to reconstruction differences. An instructive detail: interpolation obtains a lower mean CCC than the raw thick slice while more often exceeding the 0.85 threshold — it shifts features in inconsistent directions rather than improving them.
The reader study. Sixty EC-CT exams (30 healthy, 30 pneumonia), independently reviewed by two radiologists with 10 and 3 years of experience, with the original 5 mm series and the synthesised 1 mm series displayed side by side and the real 1 mm series withheld. Across 120 ratings, both readers reported added information in 100% of cases, with a mean usefulness score of 4.83 out of 5 and a mean confidence-support score of 4.69 out of 5. No rating below 4 was given.
Clinical translation. It is short, and that is the point. Out of 1,000 CT exams reconstructed at 5 mm, this work cannot say how many additional nodules would be detected, how many diameter measurements would change Fleischner category, or how many false positives would be created. None of those endpoints were measured. What the paper establishes is that automated measurements derived from a synthetic series resemble those from a real thin series more than those from interpolation — a consistency property, not diagnostic performance. One operational argument does deserve note: synthesising a whole volume takes under a minute on an 8 GB NVIDIA Quadro RTX 4000, hardware an imaging department already owns.
What is good
The dataset is the real contribution, and it is published. Two thousand exams whose two series come from the same raw projections, across two synthesis ratios and two different vendors, is exactly what the field lacked. The authors release MSS-CT and the BasicCSS code on GitHub, as de-identified NIfTI volumes, with the splits they used. The earlier resource, RPLHR-CT, was limited to a single ratio; here it serves as an external set, which is the correct use.
The pseudo-pair experiment is a well-built negative result. Six methods, three training regimes, everything else held constant — same splits, same architecture, same loss, same optimiser, same schedule — and evaluation always performed under real conditions. The ranking flip is not an anecdote: it demonstrates that a widespread evaluation protocol produces conclusions that do not survive contact with reality. Publishing this when one could have simply announced a new state of the art is an editorial choice worth crediting.
Evaluation does not stop at PSNR. Three further layers: automated segmentation across ten structures and four datasets, radiomics with a pre-declared reproducibility threshold, and human review. The external-centre adaptation analysis — a model pretrained on MSS-CT then fine-tuned on an external centre beats local training from scratch — answers a real practical question. And the authors themselves note that BasicCSS, tested without fine-tuning on RPLHR-CT, performs about as well as that set's original model, without turning it into a superiority claim.
What is less good
None of the metrics tests what is frightening about a generated image — a misleading-metric problem. PSNR, SSIM and Dice on the heart, the aorta or the ribs measure average fidelity on large structures. The risk specific to a generative model lies elsewhere: erasing a small lesion because it is improbable, or inventing a plausible one. The protocol contains no lesion-level endpoint — radiomic regions of interest are sampled at random in the lung, not centred on nodules — and the authors acknowledge that lung vessels, the smallest tubular structures evaluated, remain incompletely recovered. This is precisely the question raised by our piece on FLAIR super-resolution erasing or hallucinating small lesions, and it stays open here.
The cohorts are young and mostly healthy, whereas thin slices mainly serve older patients — population bias. MSS-CT 5 mm, the subset carrying all the external validation, has a median age of 34 and two thirds healthy subjects; EC-CT, a median of 25. Yet the prime indication for thin slices is characterising nodules, emphysema or interstitial lung disease in people over fifty. The one subset with a relevant age, MSS-CT 3 mm (median 64.5), is also the one whose clinical status is not recorded and which has no external validation at all. All centres are Chinese, only two vendors are represented, and the authors note that RPLHR-CT metadata — dose, reconstruction kernel, acquisition parameters — are missing, preventing them from explaining differences between external sets.
The reader study is designed so that it cannot fail — biased comparator, plus industry conflicts of interest. Only two readers, one with three years of experience; no interpolated comparator; no diagnostic performance endpoint; and above all, readers knew which series was synthetic. Predictable outcome: 100% added information and not a single rating below 4 across 120 evaluations, a ceiling effect that discriminates nothing. The authors state this plainly among their limitations, to their credit, but 4.83 out of 5 is the number that will circulate. Add that three of the thirteen authors are tied to Infervision Medical Technology, a commercial imaging-software vendor — one intern and two employees — declared, but weighing on work whose natural by-product is a sellable feature.
What it changes
For the research community, the operational message is direct: a ranking of synthesis methods established on simulated thick slices cannot be treated as valid under real conditions, since the winner changes. Reviewers in this subfield now have a public dataset with which to demand real-paired evaluation, and the same reasoning applies to any task where the degraded input is manufactured rather than observed. The question this paper leaves open should be the next one: build a test set enriched in annotated small lesions and measure, no longer average similarity, but erasure and hallucination rates.
For clinicians, nothing changes in practice today, and the authors do not claim otherwise: they speak of an auxiliary view, displayed next to the original thick series, never of replacing a real thin series. That framing is the right one and must hold. The reading grid worth keeping: faced with an announcement of "AI-generated thin slices", two questions suffice — was the model trained on real pairs reconstructed from the same raw data, and was anything measured beyond image similarity? If the answer to the second is no, diagnostic performance is unknown, not good.
For patients and the public, one distinction is worth keeping. A CT image reconstructed from projections is a measurement. A thin slice generated by a model from a thick slice is an inference: the detail displayed is what the model judges most probable given what it has learned, not what was observed in this patient. That can help a radiologist who only has the thick series. It does not mean the detail is real, which is why such an image must be read alongside the original, never in its place.
Further reading
The paper: Benchmarking AI-generated thin-slice CT under clinical reconstruction conditions: a multicohort study, Pengxin Yu, Haoyue Zhang, Xiaoyan Yang, Chenghao Piao, Shuiqing Zhao, Mei Xie, Xuwen Cheng, Dawei Wang, Wei Qian, Stephanie Harmon, Baris Turkbey, Yudong Wu and Shouliang Qi, npj Digital Medicine, received 2 June 2026, accepted 3 September 2026, published 16 September 2026, DOI 10.1038/s41746-026-03253-6. Open access under CC BY-NC-ND 4.0, published as an accepted version ahead of final formatting.
Affiliations: College of Medicine and Biological Information Engineering and Key Laboratory of Intelligent Computing in Medical Image, Northeastern University, Shenyang; Molecular Imaging Branch, National Cancer Institute (NIH), Bethesda; General Hospital of Ningxia Medical University, Yinchuan; Second Affiliated Hospital of Shenyang Medical College; Shenyang Medical College; Infervision Medical Technology Co. Ltd, Beijing.
Data and code: github.com/smilenaxx/MSS-CT for the MSS-CT dataset (de-identified NIfTI volumes only) and the BasicCSS implementation, in Python 3.9.23 and PyTorch 2.2. The public external set RPLHR-CT is deposited on Zenodo. EC-CT is not public: access on request to the corresponding author, subject to institutional review board approval and a data-use agreement, and restricted to non-commercial academic research. Downstream segmentation by TotalSegmentator v2.11 at default settings.
Ethics and funding: approvals from the institutional review boards of the General Hospital of Ningxia Medical University (KYLL-2025-1831), the Second Affiliated Hospital of Shenyang Medical College (2025-SYEYLL-034) and the Chinese PLA General Hospital (S2023-498-01), with informed consent waived for a retrospective study on de-identified data. Funding: National Natural Science Foundation of China (grants 82472076 and 62271131), Fundamental Research Funds for the Central Universities (N25BJD013), Department of Education of Liaoning Province (2024-LJ-10164018). Declared conflicts of interest: P.Y. is an intern at Infervision Medical Technology, X.C. and D.W. are employed there; the other authors declare none.
Also on Tatakoto: synthetic degradation of low-dose CT and its effects on nodule radiomics, and FLAIR super-resolution between erasing and hallucinating small lesions.