Pediatric chest X-ray: the AUROC survives the border crossing, the decision threshold does not
A researcher at American International University-Bangladesh froze a three-seed DenseNet121 ensemble trained on the 5,824 pediatric chest radiographs of Guangzhou — internal AUROC 0.976, sensitivity 95.1% — then applied it without any adjustment to 3,257 images from Bangladesh and 1,077 never-touched images from Vietnam. AUROC drops to 0.798 and 0.742, a severe degradation that nonetheless leaves the ranking usable; at the decision threshold frozen on the source data, sensitivity collapses to 6.2% and then to 0%, across 2,377 and 170 pneumonias respectively. Recalibrating with 163 local labels restores sensitivity to 88.3% without touching the ranking, but alerts on 78.5% of radiographs: what the paper really shows is that "it can be fixed" and "it is deployable" are not the same sentence.
The context
Pediatric pneumonia classification on chest radiographs is one of the most saturated areas of medical AI, and one of the worst evaluated. Nearly every published work since 2018 trains and tests on the same resource: the Guangzhou set released by Kermany et al. in Cell, roughly 5,800 images of children from a single Chinese centre. Systematic reviews of pediatric radiology have been repeating the same point for four years — external testing remains rare and clinical generalizability uncertain.
The problem this preprint attacks, however, is not the absence of external validation. It is its collapse into a single number. When a paper announces an "external validation", it almost always reports an AUROC and nothing else. Yet three distinct questions hide behind a performance figure. AUROC (area under the ROC curve) measures the ability to rank: draw one sick child and one healthy child at random, what is the probability the model gives the higher score to the sick one? It ignores the scale of the scores entirely. Calibration measures whether a probability announced at 0.8 really corresponds to 80% truly positive cases. The operating point, finally, is the concrete threshold above which the software raises an alert: it alone produces the reported sensitivities and specificities.
These three things move independently. A monotone shift of scores in the target population — all scores pulled downward, but in the same order — leaves AUROC intact and makes the source threshold entirely non-functional. Conversely, recalibration can repair probabilities without improving the ranking by one thousandth. That is what this work sets out to demonstrate experimentally, with a protocol locked in advance.
The method
Three cohorts, three roles. Guangzhou/Kermany (China) serves development: training, tuning, temperature scaling, threshold selection, internal testing. BDCXR-3257 (Bangladesh, 3,257 images, 880 normal and 2,377 pneumonia, a 73% prevalence) is the primary external test. VinDr-PCXR/PediCXR (Vietnam) provides a second external cohort, harmonized before any inference to 1,077 examinations (907 normal, 170 "pneumonia family"), and never used for anything other than final inference.
The analysis hierarchy is locked. This is the most important design decision. BDCXR is first evaluated as a complete cohort with the frozen source system, and that result is retained as is. Only afterwards is BDCXR cut into a 651-image adaptation pool and a disjoint 2,606-image hold-out. VinDr is never used for training, calibration, model selection or threshold selection. In other words, the author explicitly forbids himself from repainting as "external validation" a result obtained after looking at the target data.
Leakage control at the source. Every Guangzhou image is decoded and hashed: 32 exact duplicates are removed, leaving 5,824 unique radiographs. Because numeric identifiers are reused across bacterial and viral filenames, filename-derived grouping is performed inside each disease namespace. Final split: 4,165 training, 1,041 tuning, 618 internal test, with zero hash and zero derived-group overlap. The author honestly notes that these groups are leakage-control proxies, not verified patient identifiers.
The model. An ImageNet-pretrained DenseNet121 encoder, shared across two views of the same radiograph: the full image, and a lung-conditioned image produced by the TorchXRayVision anatomical PSPNet segmenter. A learned feature-wise gate combines the two representations. MixStyle regularization — which, during training, mixes style statistics across examples to discourage the model from latching onto the appearance of a particular scanner — is active on the source only. Preprocessing: robust intensity scaling, aspect-ratio-preserving letterboxing, 320 × 320, three-channel replication, ImageNet normalization; VinDr DICOMs are rescaled and MONOCHROME1 inversion handled upstream. Three seeds (42, 52, 62) are averaged at the logit level, then temperature scaling — a single parameter that flattens or sharpens probabilities — is fit on source tuning logits only (T = 0.8313).
The threshold. This is the study's most debatable choice, and it must be remembered: the primary operating point is the highest-specificity ROC point achieving at least 90% sensitivity on the source tuning data. It equals 0.9999728. Seven nines after the decimal point. A second, liberal threshold is derived post hoc but still from source data only: 0.314311.
The robustness tests. Three seed-42 matched variants on the same source split test whether the collapse depends on the proposed architecture: a conventional full-image DenseNet121, the dual view without gate or MixStyle, and the complete model. Each receives its own temperature and its own source threshold. Stress tests hunt for a shortcut signal by re-evaluating BDCXR on four views: full image, lungs, background only, border only. Separate domain classifiers measure source/target separability. Statistics: Wilson intervals for proportions, 2,000 stratified image-level bootstrap replicates for VinDr, 2,000 paired bootstraps for the fine-tuning comparison. Reporting follows CLAIM 2024 and TRIPOD+AI principles.
The results
Internally, everything is fine. AUROC 0.976, AUPRC 0.982, sensitivity 95.1% (368/387; 95% CI 92.5–96.8), specificity 90.9% (210/231; 86.5–94.0), balanced accuracy 93.0%. The kind of result that makes a headline.
In Bangladesh, the ranking holds, the decision does not. AUROC 0.798, AUPRC 0.912. But only 147 of the 2,377 pneumonias cross the source threshold: sensitivity 6.2% (CI 5.3–7.2), specificity 99.9% (879/880), balanced accuracy 53.0%. Clinical translation: per 100 radiographs passed through the system, 4.5 raise an alert and 68.5 pneumonias are missed. Calibration collapses too (ECE 0.258, Brier 0.268).
In Vietnam, the threshold never fires. AUROC 0.742 (bootstrap CI 0.701–0.782), AUPRC 0.396. None of the 170 pneumonia-family examinations crosses the threshold: sensitivity 0% (Wilson CI 0–2.2), specificity 100%, balanced accuracy 50.0%. The calibration slope is 0.196. Two narrower endpoint definitions give AUROCs of 0.729 and 0.734, and the same zero sensitivity.
This is not the proposed architecture's fault. The conventional full-image DenseNet121 goes from 0.961 to 0.749 on BDCXR, with sensitivity from 94.8% to 16.2%. Dual view without gate: 0.977 → 0.766, sensitivity 96.1% → 4.7%. The complete model: 0.966 → 0.789, 95.9% → 4.4%. The phenomenon is identical throughout; the "improved" model gains a little external AUROC and loses more threshold sensitivity.
Recalibration repairs the probabilities, not the decision. On the 2,606-image hold-out, the locked source model gives AUROC 0.799, sensitivity 6.4%, ECE 0.256. Platt recalibration — a simple logistic transform of the score, therefore monotone, therefore with no effect on AUROC — fit on 163 local labels (5% of the cohort) gives: sensitivity 88.3%, ECE 0.044, but specificity 47.9%, PPV 0.821, alert rate 78.5%, and 14.1 false alerts per 100 examinations. With 326 labels: sensitivity 91.9%, specificity 38.6%, alert rate 83.7%, 16.6 false alerts per 100. With 651 labels, no further improvement (90.7% / 42.5%). The label budget saturates immediately, and what it buys is a policy that flags four radiographs out of five.
And that repair is unstable. Across 200 repeated draws of 163 labels, mean sensitivity is 91.5% (2.5th–97.5th percentiles: 86.8–95.5%) but mean specificity is 37.1% with a range of 21.0 to 49.3%. The direction of recovery is reproducible; the operating point obtained depends heavily on which 163 images were labelled.
Fine-tuning adds nothing. The prespecified 326-label comparison — last-block fine-tuning versus matched Platt recalibration — gives 0.7907 against 0.7893, that is +0.00136 (95% CI −0.00117 to +0.00382; p = 0.293). Three-seed ensembles reach 0.805 with 326 labels and 0.815 with 651: real but marginal ranking gains.
The signal exists outside the lungs. On BDCXR, pneumonia AUROC stays above chance on the full image (0.794), on the lungs (0.760), on the background only (0.719) and on the border only (0.667). Source-versus-target classification is near perfect on every view.
The pathological threshold explains only part of the disaster. The liberal threshold (0.314311) raises internal sensitivity to 100%, but yields only 69.0% in Bangladesh and 15.9% in Vietnam. The seven-nines threshold amplified the failure; it did not create it.
What is good
The analysis hierarchy is locked and the author holds to it to the point of self-incrimination. BDCXR is evaluated as a complete cohort and that result is retained before a single local label is looked at; VinDr is never touched. This is exactly the discipline TRIPOD+AI asks for and almost nobody applies, because it forbids presenting as external a result obtained after adaptation. The author goes further: he discloses that an earlier version of his own pipeline compared raw institution name strings and leaked, and that the released folds were re-cut. He also states that the Guangzhou leakage-control groups are filename-derived and are not verified patient identifiers. This level of self-reporting is rare.
The architecture-robustness arm defuses the most obvious objection. A hurried reader would conclude that the gated dual view is a bad idea. The author tests precisely that: a bare DenseNet121, trained on the same split with its own calibration and its own threshold, collapses too (0.961 → 0.749, sensitivity 94.8% → 16.2%). The transport result is therefore not an artefact of one implementation. This is a negative contribution about his own architecture, published as such: the author explicitly writes that the paper does not present a novel vision backbone and that the contribution is the reproducible decomposition of the failure.
"Recovery" is costed in workload, not only in metrics. This is what separates the paper from the domain-adaptation literature. Almost every work showing that a light recalibration restores sensitivity stops at sensitivity. Here one reads, on the same table row: sensitivity 88.3%, specificity 47.9%, alert rate 78.5%, 14.1 false alerts per 100 examinations. And the 200-resample analysis shows that the specificity obtained varies from 21% to 49% depending on the draw. In other words, the author measures not only the operational cost of the repair, but also its unpredictability.
What is less good
The primary threshold at 0.9999728 manufactures part of the result the abstract leads with. An operating point at seven nines is not a clinical choice, it is the symptom of a saturated sigmoid on an easy source set — the model outputs probabilities glued to 1 for almost every Guangzhou pneumonia, so the rule "highest specificity at sensitivity ≥ 90%" lands on an absurd point. The headline figures — 6.2% and 0% — are partly the product of that mechanism. The author acknowledges it, tests a liberal threshold, and rightly concludes that it "amplified but did not create" the failure. But the abstract still leads with 6.2% and 0%, whereas the honestly transportable figures are closer to 69.0% in Bangladesh and 15.9% in Vietnam. This is a textbook misleading metric — here to the model's detriment rather than its advantage, which is still a distortion.
There is no human comparator, and the reference standard is weak on all three floors. Nothing is compared to a radiologist, a pediatrician, or even a simple clinical rule: we therefore cannot tell whether 69% sensitivity in Bangladesh is good, bad or beside the point. The Guangzhou labels come from a resource whose limitations have been documented since 2018. The Vietnamese endpoint had to be harmonized by the author himself before inference (Pneumonia, Broncho-pneumonia and Pleuro-pneumonia counted positive, clean "No finding" counted negative) — a defensible choice, frozen before analysis and checked against two narrower definitions, but still an experimenter decision about the variable being measured. Finally, BDCXR has no verified patient identifier: the image-level bootstrap cannot account for possible multiple radiographs of the same child, which probably understates uncertainty — the author says so in black and white.
The title says "three countries" while the study can attribute nothing to country. The three resources differ simultaneously in geography, age distribution, prevalence (73% versus 15.8%!), equipment, compression, disease spectrum, annotation protocol and reference standard. The author says so explicitly in the limitations and refuses any causal conclusion about geography — but the title and abstract keep selling "across three countries". Three reproducibility weaknesses follow. The shortcut analysis is suggestive, not causal: background predictive at 0.719 shows that label-correlated signal exists outside the anatomy, not that the model used it, and the author concedes this. The code is announced as a "reproducibility archive supplied with this submission" — no public repository, no code DOI, no released weights, and the preprint carries a CC BY-NC-ND 4.0 licence, so neither commercial nor derivative use. And this is single-author work, unfunded, with no prospective validation, no local radiologist re-adjudication and no decision curve: what the paper rigorously documents is a transport failure, not a path to deployment.
What it changes
For the research community. The transferable result is none of the numbers, it is the five-component protocol: discrimination, calibration, frozen operating-point behaviour, shortcut signal, and recoverability under a limited label budget. Those five axes cost a few extra compute hours on data one already owns, and they separate situations that AUROC alone conflates completely. The paper's sharpest demonstration fits in one line: AUROC 0.742 and sensitivity 0% on the same cohort, with the same model, at the same instant. Any systematic review counting "successful external validations" on the basis of an AUROC is counting wrong. We had already seen this split at work in the transportability of ICU delirium models between eICU and MIMIC, and the mechanism is identical.
For clinicians and device buyers. The number to remember is 78.5%. That is the alert rate after a recalibration presented as successful — sensitivity restored to 88.3%, excellent calibration at an ECE of 0.044. Software that flags four radiographs out of five in a department producing two hundred a day is not a triage tool, it is a noise source. The practical consequence is that a contract clause requiring "an AUROC ≥ 0.80 on our data" is insufficient: one must require the threshold, the sensitivity and the specificity and the alert rate at that threshold, on the local population, plus the recalibration procedure for when they drift. The corollary is that a model bought with its factory threshold is probably unusable as delivered — but recalibrating it locally requires local labels, hence radiologist time, hence a budget nobody provisions.
For patients and the public. Across the 1,077 Vietnamese examinations, the model detected zero pneumonias out of 170. Not because it "sees" nothing — its ranking stays clearly above chance — but because the dial set in China corresponded to nothing in Vietnam. This is the most insidious failure mode of medical AI: the system does not err loudly, it goes silent. And the fact that the background and border of Bangladeshi radiographs remain predictive of pneumonia is a reminder that these models partly learn the hospital, not the disease — the same shortcut learning we described in the fusion of screening mammography datasets. The general lesson, for a non-specialist reader: when a press release announces that a model reaches 95% sensitivity, the question to ask is not "on how many patients?" but "at what threshold, and was that threshold set on the same data?".
Further reading
The preprint: Cross-dataset transportability of pediatric chest X-ray deep learning across three countries: discrimination, calibration, operating-point failure, and limited-label recovery, arXiv:2609.05140 [eess.IV], DOI 10.48550/arXiv.2609.05140, posted 4 September 2026, under a CC BY-NC-ND 4.0 licence. Single author: Nazim-E-Alam, Department of Computer Science, American International University-Bangladesh (AIUB), Dhaka. No funding from public, commercial or not-for-profit sectors; no competing interests declared. The author declares having used ChatGPT (OpenAI) for literature synthesis, code debugging, manuscript structuring and language refinement, and states that numerical claims were verified against preserved analysis outputs; no scientific image was generated or altered by AI. The source data remain with their respective custodians: Guangzhou/Kermany (Cell, 2018), PediCXR (Scientific Data, 2023) and VinDr-PCXR on PhysioNet, whose DICOMs are access-restricted and not redistributed. A reproducibility code archive accompanies the submission but is not publicly deposited. On the methodological background, two founding reads: Zech et al., PLOS Medicine 2018, on a pneumonia model encoding site identity, and Van Calster et al., BMC Medicine 2019, "Calibration: the Achilles heel of predictive analytics". The reporting checklists cited are CLAIM 2024 and TRIPOD+AI.