PhysioFusion in cardiac surgery: what fusing 26 intraoperative signals with the STS registry changes in 1,393 patients
A team from the Roux Institute (Northeastern University) and Maine Medical Center built PhysioFusion, an ensemble of models that combines preoperative Society of Thoracic Surgeons (STS) registry data with 26 intraoperative monitoring signals to predict complications after cardiac surgery, in 1,393 patients from a single hospital. The AUC reaches 0.87 in cross-validation, but the registry variables alone do just as well (0.87): the intraoperative signal adds nothing measurable here. Read it as an honest internal study with a negative finding, not as proof that multimodal fusion improves prediction.
The context
After cardiac surgery, a non-trivial share of patients suffers a major complication (11.6% in the cohort studied here): prolonged ventilation, stroke, kidney failure, unplanned reoperation, deep sternal wound infection, death. In the United States, the Society of Thoracic Surgeons (STS) runs a national registry in which every operation is coded, and its preoperative risk scores are the reference for comparing centers. These scores use only what is known before the incision.
Yet an operation lasting several hours produces a continuous stream of measurements: arterial and venous pressures, electrocardiogram, oxygen saturation, temperature, cardiopulmonary bypass times. The natural idea of this paper is that these time series carry a risk signal the preoperative variables miss, so a model that fuses them should predict better. It is a preprint posted on medRxiv in September 2026 and has not yet been peer reviewed.
The method
The data. 1,393 consecutive adult patients operated on at Maine Medical Center (MaineHealth) between September 2022 and April 2024, of whom 161 (11.6%) had at least one of the eleven adverse events endorsed by the STS and the National Quality Forum, grouped into a composite outcome. Static variables come from the STS database (age, history, cardiac function, type of surgery, bypass duration, and so on). The time series cover 26 intraoperative signals.
The models. The core is a two-branch network: a residual network (stacked blocks whose output is added to their input, which stabilizes training) for the static variables, and for the time series a convolutional network (which detects local patterns in the signal) and a bidirectional GRU (a recurrent network that reads the series in both directions). Alongside it sit four classical models: XGBoost (successive decision trees, each correcting the errors of the previous one), L1-penalized logistic regression, a support vector machine (SVM) and a two-stage gradient boosting cascade. Predictions are combined by a weighted average, then recalibrated with isotonic regression (an increasing curve that makes "the model says 20%" match "20% of patients actually have the event").
The validation. Stratified five-fold cross-validation with shuffling: the dataset is split into five parts, each serving once as the test set. Class imbalance is handled with SMOTEENN (creating synthetic positive cases and cleaning ambiguous ones), on training folds only. Three decision thresholds are evaluated: one maximizing positive predictive value at sensitivity ≥ 0.7, one maximizing F1, and a "balanced" one where sensitivity equals specificity.
The results
ML metrics. At the balanced threshold, the combined model reaches an AUC (the probability of ranking a patient with an event above a patient without one) of 0.87 (95% CI 0.82 to 0.91), sensitivity 0.76, specificity 0.83, positive predictive value (PPV) 0.40 and negative predictive value (NPV) 0.96. Calibration error is 0.036.
The central result is the breakdown by modality. Static variables alone give an AUC of 0.87 (0.82 to 0.92), sensitivity 0.74, specificity 0.82, PPV 0.45 and better calibration (0.023). Time series alone give 0.83 (0.80 to 0.86), with a PPV of 0.33. The intervals overlap almost entirely: on this cohort, fusion does no better than the registry data. A signal-by-signal ablation concludes that no single monitoring channel is indispensable, with mean arterial pressure and central venous pressure the least substitutable. Among the individual models, gradient boosting reaches 0.86, logistic regression 0.78, the SVM 0.69 and the deep network 0.72 (0.60 on time series alone).
Clinical translation. For 1,000 patients with the cohort prevalence (116 events), a sensitivity of 0.76 means about 88 complications caught and 28 missed. A specificity of 0.83 among the 884 patients without an event means about 150 false alarms. A direct calculation gives a PPV near 0.37, slightly below the reported 0.40, probably due to rounding and per-fold computation. In practice: roughly two alerts in three would be false, and a little more than one complication in four would escape the model. The NPV of 0.96 is reassuring but mostly reflects how rare the event is: predicting "no complication" for everyone would have an NPV of 0.88.
What is good
A negative result reported without dressing it up. The authors report that the two modalities are substitutable rather than presenting fusion as a success. That is rare in a field where complexity is often sold as a virtue in itself, and it is exactly the information needed by anyone considering equipping an operating room.
A carefully built evaluation on the technical points. SMOTEENN rebalancing is restricted to training folds, test folds keep the real 11.6% prevalence, three operating thresholds are reported instead of one, and confidence intervals are given. Five model families show that the result does not hinge on an architecture choice.
Consecutive hospital data and a standardized outcome. Patients are consecutive (no cherry-picking of easy cases) and the outcome follows STS definitions rather than an ad hoc one. A no-competing-interest statement appears in the text.
What is less good
Strictly internal validation (population and center bias). One hospital, 161 events, no external or temporal validation: cross-validation measures performance on the same patients, the same operating room, the same monitoring devices. The authors themselves acknowledge that the small number of events penalizes high-capacity models, which explains the poor deep network score (0.72), but it also means the comparison between architectures says little about what a larger dataset would show.
A missing comparator (biased comparator by omission). The text we could consult applies no established STS risk score. That is the comparator that matters: is a model with an AUC of 0.87 better than the score the surgeon already has at hand? The internal comparators are components of the ensemble, not clinical references. Without this, we cannot say whether the model beats current practice.
Information leakage risks and a metric to read with care (data leakage, misleading metric). The "balanced" threshold is chosen on the evaluated predictions, and the isotonic calibration is fitted on model predictions; the text we consulted does not say whether both steps are strictly separated from the test folds, which could make sensitivity, specificity and calibration optimistic. At 11.6% prevalence, an AUC of 0.87 comes with a PPV of 0.40: most alerts would be false. The composite of eleven very heterogeneous events (a death and a reoperation have neither the same mechanism nor the same prevention) also complicates any targeted clinical action. Finally, code and data are available only on request.
What it changes
For the research community. The paper provides a useful test case: on a modest single-center cohort, raw intraoperative signals add nothing to a well-filled registry. The next question is whether this holds with more patients, external validation and a better-designed signal representation (pretraining on large volumes of monitoring data, for example), rather than stacking more networks on 161 events.
For clinicians. Nothing changes in practice: the model is not validated outside the center, not compared with the STS score and not evaluated prospectively. It does remind us that the preoperative registry already holds most of the predictable signal, which argues for improving data entry quality before investing in new sensors or data streams.
For patients and the public. These models do not replace the surgical team's judgment. A tool whose alerts are right one time in three remains a triage aid for the team, not an individual verdict, and patients should not expect "smart" operating-room surveillance in the short term.
Going further
The paper (preprint, not peer reviewed): PhysioFusion: A Multi-Modal Ensemble using Static Preoperative Variables and Time Series Intraoperative Data to Predict Adverse Events Following Cardiothoracic Surgery, Rajashekar Korutla, Anne Hicks, Marko Milosevic et al. (Roux Institute, Northeastern University; Spectrum Healthcare Partners; Maine Medical Center, MaineHealth), medRxiv, September 2026, DOI 10.64898/2026.09.24.26363920. Code and data: on request from the corresponding author, under a data use agreement. The authors declare no competing interests; funding was not identified in the text we consulted.
On models trained on very few events, see our review of the early hepatocellular carcinoma recurrence study. On a single-center clinical prediction model, see our review of the Mohs surgery study.