Gram-negative bacteremia in the ICU: three frozen models, prospectively validated, do worse than a four-item bedside score
An infectious diseases and intensive care team at Şişli Hamidiye Etfal Training and Research Hospital in Istanbul took three machine learning models it had developed in-house to predict Gram-negative bacteremia in the ICU, froze them — no retraining, no threshold tuning, outputs blinded during recruitment — and tested them prospectively on 289 consecutive blood cultures. All three top out at an area under the ROC curve of 0.61, which is worse than the Pitt Bacteremia Score (0.63) and the quick Pitt score (0.64), two bedside scores that take thirty seconds and no computer. The 0.70 figure highlighted in the abstract comes from a subgroup of sixteen events, called predefined in the methods and post-hoc in the abstract, and the best model's calibration curve shows an eighth risk decile three times less often bacteremic than the seventh.
The context
In the ICU, bacteremia — live bacteria circulating in the blood — is confirmed by blood culture, and a blood culture takes twenty-four to seventy-two hours. During that window the intensivist decides alone: antibiotics or not, which ones, at what spectrum. Every hour of inappropriate therapy costs mortality, particularly for Gram-negative bacilli of the Enterobacterales family. Hence the long-standing idea of a tool that would predict the blood culture result before the blood culture.
Two families of tools already coexist, and they do different jobs. Severity scores — SOFA, SIRS, NEWS-2 — measure organ failure and the scale of the inflammatory response; they were built to stratify sepsis, not to guess a blood culture. Bacteremia-specific scores — the Pitt Bacteremia Score (PBS) and its abbreviated version, the quick Pitt (qPitt) — were built to predict mortality once bacteremia is confirmed. None of the five was designed for the task evaluated here.
Against them, the machine learning literature on sepsis and bacteremia is abundant. It has a structural flaw the authors cite themselves: only 6% to 14% of models developed on ICU data have undergone external validation, and performance almost always drops when they do. The Istanbul team had previously published three models trained retrospectively on routine data from its own ICUs. This paper is the next step, the one almost nobody takes: testing them in real conditions, untouched.
The method
A prospective design, locked models. Single-centre observational study, 1 February to 31 July 2025, in the adult ICUs of Şişli Hamidiye Etfal. Eligible were blood cultures drawn from patients aged 18 or over after at least 48 hours in the ICU and meeting the Sepsis-3 definition of suspected infection: a culture obtained within 24 hours of starting antibiotics, or antibiotics started within 72 hours of the culture. That filter matters, and we will come back to it: it restricts the analysis to episodes where a clinician actually suspected something.
323 consecutive blood cultures were collected. Excluded were 17 fungemias, 15 Gram-positive bacteremias and 2 ineligible Gram-negative species. That leaves 289 samples: 192 culture-negative and 97 Gram-negative bacteremias, a prevalence of 33.6%. The 101 isolates are dominated by Klebsiella pneumoniae (56.4%), Acinetobacter baumannii (15.8%) and Pseudomonas aeruginosa (11.9%); 74.3% are carbapenem-resistant and 39.6% colistin-resistant. The source was respiratory in 61.9% of cases.
The three models. A Random Forest — a forest of decision trees that vote — and two mixed-effects variants: BiMM (binary mixed model) and MERF (mixed-effects random forest). "Mixed effects" add a patient-specific term to the model, which handles repeated measurements in the same person properly instead of treating them as independent — relevant here, since each patient contributes a 48-hour trajectory of vital signs and laboratory values.
What makes the validation credible. The models were applied as-is: same predictors, same observation window, no retraining, no hyperparameter tuning. The decision thresholds, derived by the Youden index on the training set of the development study, were carried over without recalibration. Model outputs were not accessible during recruitment, and the analyses were run only once data collection had closed. That is the definition of a clean temporal validation.
The comparators. ΔSOFA ≥ 2, SIRS ≥ 2, NEWS-2 ≥ 5, PBS ≥ 6, qPitt ≥ 3, plus the Sepsis-3 diagnosis, all at predefined cut-offs. Data came from the 48 hours preceding sampling, recorded at 10 a.m. to limit intra-day fluctuation, in complete-case analysis. Analyses in R 4.4.1, reported against the TRIPOD and STARD checklists. The study is registered on ClinicalTrials.gov as NCT07126106 — on 17 August 2025, seventeen days after data collection ended.
The results
In the full cohort, the models lose. AUROC of 0.61 for all three models, against 0.64 for qPitt, 0.63 for PBS and for NEWS-2, 0.62 for ΔSOFA. Balanced accuracy — the mean of sensitivity and specificity, which corrects for class imbalance — runs from 60% to 61% for the models. Specificities are low, 43% to 50%, and positive predictive values sit around 41-42%. BiMM achieves the best sensitivity among the models (78%), but PBS reaches 93%.
In the subgroup, everyone improves, and the scores improve faster. Restricted to blood cultures drawn between ICU days 4 and 10, the analysis covers 89 samples with 16 bacteremias. Model AUROCs rise to 0.69-0.70 and balanced accuracy to 68%. But PBS reaches 0.77 and qPitt 0.76. PBS at a cut-off of 6 catches 100% of bacteremias; BiMM and MERF reach 90% sensitivity; Random Forest takes the other end of the trade-off, with 74% specificity for 63% sensitivity.
MERF's calibration is broken. The authors ranked patients into deciles of predicted risk and counted actual bacteremias in each. The rate rises steadily across the first seven deciles, then drops from 58.8% in the seventh decile to 17.2% in the eighth. It climbs again afterwards without ever regaining the seventh's level. In other words, the patients the model flags as highest risk are less often bacteremic than those it places mid-table.
Without prior antibiotics, the models predict nothing. An exploratory stratified analysis splits the cohort by exposure to a Gram-negative-active antibiotic before sampling. Among the 65 unexposed patients, the three models' AUCs run from 0.49 to 0.53 — chance. Among the 224 exposed patients, they all rise to 0.65. The development cohort had roughly 70% of patients already on antibiotics.
Clinical translation. Take MERF in the main cohort, over 1,000 blood cultures at 33.6% prevalence. The model would flag about 620 patients; 255 would indeed have Gram-negative bacteremia, 365 would not. It would miss 81. Three alerts in five are false, and one bacteremia in four slips through. In the subgroup, where prevalence falls to 18%, 90% sensitivity costs more still: per 1,000 samples, roughly 605 alerts for 162 real bacteremias, or close to three false alerts for every true one. PBS at 100% sensitivity does worse on this front: its 32% specificity yields a positive predictive value of 24%. In a unit where 74% of isolates resist carbapenems, every false positive is a potential escalation to colistin or last-line beta-lactams.
A useful negative result, tucked into Table 1. Neither CRP (181 ± 92 vs 175 ± 85 mg/L, p = 0.6), nor procalcitonin (11 ± 27 vs 7 ± 22 ng/mL, p = 0.2), nor temperature (37.2 °C in both groups, p = 0.4), nor white cell count (p = 0.070) separates bacteremic patients from the rest. What does separate them: age (71 ± 17 vs 67 ± 16 years, p = 0.046), invasive mechanical ventilation (84% vs 69%), central venous catheter (76% vs 64%), vasopressors, and above all isolation of a Gram-negative organism in the preceding thirty days (57% vs 34%, p < 0.001). The patient's microbiological history beats the inflammatory markers.
What is good
The models were genuinely frozen, and that is rare. No retraining, no post-hoc search for an optimal threshold, the development study's Youden cut-offs carried over unchanged, and outputs blinded throughout collection. The opposite temptation — re-thresholding on the validation cohort and then calling the result "validation" — is the field's most widespread sin, and it mechanically inflates performance. Here the protocol forbids that breathing room. That is exactly what makes the 0.61 figure believable, and it is the same problem raised by frozen thresholds carried from one country to another.
The comparator is honest, and it wins. The authors could have pitted their models against SIRS alone, the weakest of the five scores (AUROC 0.57), and declared superiority. They put PBS and qPitt in the same table, watched them beat their own models in both cohorts, and wrote it down. That row of Table 3 is the paper's most informative result, and it works against its authors.
The validation population is defensibly defined. By restricting the analysis to episodes meeting the Sepsis-3 definition of suspected infection, the authors exclude patients nobody suspected of bacteremia — precisely the ones whose inclusion artificially inflates specificity in many earlier studies, which they name explicitly. Add the publication of the decile calibration curve, which contradicts their own model, and completed TRIPOD and STARD checklists.
What is less good
The headline result rests on sixteen events, and its status changes from page to page. The day 4-10 subgroup holds 89 samples and 16 bacteremias. With sixteen events, uncertainty around an AUROC is considerable — and no confidence interval accompanies any value in Table 3, for models or scores. More awkwardly: the Methods section describes this subgroup as "predefined", the abstract calls it "post-hoc". Both cannot be true, and registering the study on ClinicalTrials.gov on 17 August 2025, seventeen days after collection closed, settles nothing — a retrospective registration protects against none of the biases registration is meant to prevent. This is the classic misleading metric failure mode: the only figure in the abstract belongs to the most favourable and least populated stratum.
The outcome label is contaminated by treatment. 78% of patients were already receiving a Gram-negative-active antibiotic when sampled, and that exposure was more frequent among the culture-negative (82% vs 69%, p = 0.022). Some of the "negatives" are therefore likely bacteremias sterilised by antibiotics: the model is being scored against a noisy reference standard, exactly the problem documented elsewhere for labels extracted from clinical routine. The authors acknowledge it, but their reading of the stratified analysis deserves inverting: they conclude the models "suit" already-treated patients better. The same data reads another way — among the 65 unexposed patients, those whose label is most reliable, AUCs of 0.49 to 0.53 mean the models have no discriminative power at all. The apparent signal exists only where the reference standard is most doubtful. Whether the models capture an infection or a prescribing trajectory is unresolved.
Calibration invalidates the main use one would make of the model. A drop from 58.8% to 17.2% between the seventh and eighth deciles is not imprecision, it is an inversion. A model exists precisely to rank patients by descending risk so as to decide whom to treat first; if its top decile is less often ill than its middle decile, that ranking cannot carry a decision. The authors invoke variables with disproportionate weight and suggest the mid-range might still be usable. That hypothesis is plausible, but it is not tested, and no standard calibration measure — slope, intercept, Brier score — is reported. The same blind spot has already been noted for alert policies in intensive care. Add the size constraints: a single centre, no a priori sample size calculation, complete-case analysis, and a local epidemiology — 74% carbapenem resistance — transposable almost nowhere in Western Europe.
What this changes
For the research community, this paper is useful for what it fails to find. It documents, with a protocol few teams impose on themselves, the gap between development performance and real-world performance on locked models. The comparison the authors draw themselves is illuminating: Bhavani and colleagues reported an AUROC of 0.78 across 252,569 hospitalised patients, Ming and colleagues an AUROC of 0.97 across 20,850 patients — but in the latter study's hospital-acquired bacteremia subgroup, AUROC fell back to 0.64, close to Istanbul's 0.61. ICU-acquired bacteremia, in patients already on antibiotics and already colonised with resistant flora, appears to be an intrinsically harder problem than bacteremia on admission. The avenue this work opens is methodological rather than algorithmic: as long as the label "negative blood culture" depends on the treatment received, no architecture will fix the problem, and it will have to be handled explicitly.
For clinicians, the practical conclusion is uncomfortable for the field but clear. In this cohort, the Pitt Bacteremia Score and the quick Pitt — five and four items, computable in your head at the bedside, free, requiring no information system — do better than three machine learning models fed forty-eight hours of vital signs and laboratory values. Nothing here justifies deploying an algorithmic tool for this task today. And none of the nine tools tested, model or score, has a positive predictive value above 42%: a positive signal is never on its own enough to trigger escalation, especially in a resistance setting where escalation means colistin or last-line beta-lactams. The directly transferable bedside finding is not a model at all: it is that isolation of a Gram-negative organism in the preceding thirty days discriminates better than CRP, procalcitonin, temperature and white cell count, which do not discriminate at all.
For patients and the public, two orders of magnitude are worth keeping. An AUROC of 0.70, often called "moderate" in papers, does not mean the model is right seven times out of ten: it means that if you draw at random one bacteremic patient and one who is not, the model assigns the higher risk to the right one 70% of the time. And a positive alert, here, is wrong roughly three times out of four. That is the ordinary price of a tool tuned to miss nothing in a serious disease — a price paid in broad-spectrum antibiotics, side effects, and selection pressure on resistant bacteria.
Further reading
The paper: Prospective validation of machine learning models predicting Gram negative bacteremia in ICU versus clinical scores, npj Digital Medicine, published 14 September 2026, DOI 10.1038/s41746-026-03239-4, open access under CC BY-NC-ND 4.0. Authors: Ahmet Doğukan Bayrak, Olcay Dilken, Gizem Çamkerten, Hakkı Meriç Türkkan, Ceren Atasoy Tahtasakal, Mustafa Altınay, Dilek Yıldız Sevgi, İlyas Dökmetaş and Okan Derin — infectious diseases and intensive care departments at Şişli Hamidiye Etfal Hospital (Istanbul), department of statistics at Yıldız Technical University, adult intensive care at Erasmus MC Rotterdam, and the epidemiology doctorate programme at Istanbul Medipol University.
Registration: ClinicalTrials.gov NCT07126106, filed 17 August 2025. Ethics approval no. 2880 from the Şişli Hamidiye Etfal hospital committee, 14 January 2025, with written consent from patients or their legal representatives. TRIPOD and STARD checklists provided as supplementary material.
Data and code: neither the de-identified datasets nor the analytical code is deposited in a public repository; both are stated to be available from the corresponding author on reasonable request. Analyses in R 4.4.1. The authors declare no external funding and no competing interests.
The development study for the three models, by Dilken and colleagues, covers short clinical and laboratory trajectories in intensive care. For the comparators cited: the Pitt Bacteremia Score (Al-Hasan and Baddour, Clinical Infectious Diseases, 2020), the quick Pitt score (Battle and colleagues, Infection, 2019), and the Sepsis-3 definitions (Singer and colleagues, JAMA, 2016).