SIMPLE-HF, an 11-variable mortality score for heart failure: what a Transformer distilled from 373,000 UK patients changes
A team at the University of Oxford built SIMPLE-HF, a mortality score for heart failure that condenses a Transformer model trained on the electronic health records of 373,389 UK patients into just 11 variables measurable in routine care, with no echocardiogram required. On a validation cohort of 77,712 patients, the score identifies twice as many actual deaths as MAGGIC-EHR — a version of the reference MAGGIC score adapted for routine data — for the same number of patients flagged as high risk. The work is methodologically solid, but remains an internal validation within a single healthcare system, against a comparator deliberately weakened relative to the original MAGGIC score.
The context
Heart failure affects roughly 64 million people worldwide and remains one of the deadliest chronic conditions in internal medicine: a large share of diagnosed patients die within five years. Stratifying that risk — knowing which patients, whether followed in primary care or discharged from hospital, most need closer follow-up, treatment escalation, or a specialist opinion — directly shapes how cardiology resources are allocated.
The most widely used tool for this, MAGGIC (Meta-Analysis Global Group in Chronic Heart Failure), was built in 2013 from a meta-analysis of thirteen cohorts. It combines simple clinical variables (age, blood pressure, heart rate) with data that require a specialist test: NYHA functional class (breathlessness on exertion) and left ventricular ejection fraction (LVEF), measured by echocardiography. The Oxford team documents a two-part problem. On one hand, MAGGIC discriminates poorly — its ability to tell a high-risk patient from a low-risk one remains modest. On the other, its key variables are often missing from routine electronic health records (EHR), simply because no recent echocardiogram was performed. More complex AI models, trained directly on longitudinal EHR sequences, are more accurate but remain black boxes that are hard to integrate into a care pathway and hard to get clinicians to trust.
The paper's goal is to escape this dilemma: keep the accuracy of a complex model trained on longitudinal data, without keeping its complexity of use.
The method
The team worked with CPRD Aurum (Clinical Practice Research Datalink), which aggregates general-practice records for a representative slice of the English population. Out of 19,094,480 individuals, 373,389 adults with an incident heart failure diagnosis between 2010 and 2020 were retained. A notable methodological point: the split between training data (1,153 practices, 295,677 patients) and validation data (289 practices, 77,712 patients) was made at the level of GP practices, not individual patients — to prevent a single practitioner's coding habits from appearing in both the training and test sets, which would have artificially inflated the measured performance (a classic case of data leakage).
Building the final model, SIMPLE-HF (Simplified Intelligent Mortality Prediction for Longitudinal EHRs), followed three steps. First, a baseline model — a multi-layer perceptron (MLP), a simple multi-layer neural network — was trained within a survival framework (SODEN, which models not a binary outcome but the time until death) using the MAGGIC score's variables. Next, a separate Transformer model — the architecture behind large language models, here applied to a patient's temporal sequence of diagnoses and prescriptions — was trained on the cohort's full longitudinal data to identify which comorbidities carried the most prognostic information. Finally, SHAP (SHapley Additive exPlanations, which assigns each variable a quantified contribution to the prediction, borrowed from cooperative game theory) was used to distill that set of 21 candidate variables down to a final model with just 11: year of birth, age, body mass index, how long ago heart failure was diagnosed, prescription of beta-blockers or ACE inhibitors/ARBs, and a history of cancer, acute kidney injury, pneumonia, cardiac arrest, respiratory failure, and chronic obstructive pulmonary disease (COPD).
The Transformer is therefore not the final model: it is a feature-discovery tool, whose findings are then transferred into a simple score computable from data present in any general-practice record — no echocardiogram needed. For comparison, the authors had to build MAGGIC-EHR, a version of the original MAGGIC score adapted to routinely available data, since the missing NYHA class and LVEF would otherwise have made comparison impossible for a large part of the cohort.
The results
On the validation cohort, SIMPLE-HF reaches a C-index of 0.801 (95% CI: 0.795–0.806) versus 0.735 (0.728–0.741) for MAGGIC-EHR. The C-index measures the probability that, for two randomly chosen patients, the model correctly ranks which one will die first — 0.5 corresponds to a random guess, 1.0 to a perfect ranking. On AUPRC (area under the precision-recall curve, more informative than standard AUC when events are frequent — here roughly 42% five-year mortality), the gap is similar: 0.684 versus 0.560. Both models remain well calibrated — their calibration curves, smoothed via spline regression, stay close to the ideal reference.
Clinical translation: at a 60% risk threshold, out of 1,000 patients screened, SIMPLE-HF flags 257 as high risk, and 199 of them will actually die, versus 145 patients flagged and only 99 deaths among them for MAGGIC-EHR. In other words, for a comparable clinical workload (a little under 260 patients to reassess per 1,000), SIMPLE-HF captures twice as many actual deaths. At a 50% threshold for 12-month mortality, the score identifies 4,218 additional true positives and reduces missed deaths by 4,818, for a gain of 7.2 points of positive predictive value. Decision curve analysis, which measures a decision tool's net benefit across clinical thresholds, confirms SIMPLE-HF's advantage across the entire range of thresholds tested, up to around 60%.
What is good
A validation scale rare in this field. 373,389 patients followed over a decade (2010–2020), drawn from a database covering a representative slice of the English population — an order of magnitude above most published cardiovascular risk scores, many of which rely on cohorts of a few thousand patients.
A train/validation split that takes data leakage seriously. Splitting by GP practice rather than by individual patient is a sound and rarely-this-explicitly-documented methodological precaution: it prevents a given practice's coding style from appearing in both training and test sets.
A complete clinical double-display. Beyond the C-index, the authors provide a decision curve analysis and a translation into true positives and false negatives gained per 1,000 patients screened — exactly the kind of conversion that makes a score usable by a clinician rather than an abstract performance number. The final score requires no specialist test, making it a potentially deployable tool in primary care.
What is less good
No external validation — the most classic failure mode of AI-health risk scores. Despite its scale, the entire validation remains internal to a single database, a single country, a single healthcare system (the English NHS). The authors acknowledge this explicitly in their limitations: "further external validation in diverse international cohorts and healthcare systems is necessary to confirm generalisability of our approach in other settings." Splitting by GP practice limits the risk of data leakage but guarantees nothing about transportability to a different coding system (US hospital records, for instance) or to a population with different comorbidities.
A comparator weakened by design. MAGGIC-EHR is not the original MAGGIC score: it is an adaptation from which the authors had to remove major predictive variables — NYHA class and LVEF — because they could not find them in routine data. Part of the measured performance gap (0.801 versus 0.735) therefore reflects both SIMPLE-HF's real superiority and the handicap imposed on the comparator to make the comparison possible at all. The authors disclose this honestly, but a reader who only remembers the two C-index numbers risks overestimating the actual gap with the best possible use of MAGGIC, with echocardiography.
Limited reproducibility and non-trivial industry ties. CPRD data is not freely accessible — its use requires a license and approval from an independent scientific committee — and nothing indicates that the Transformer's or SIMPLE-HF's code or weights have been released. Independent replication is therefore currently impossible. In addition, the funding (a Horizon Europe grant) comes alongside declared industry ties for several authors: lead author Kazem Rahimi discloses consulting fees from Medtronic and Lucem Health, as well as research funding from the Novo Nordisk Foundation and Roche; another author also discloses consulting fees from Lucem Health. This does not disqualify the results, but is worth knowing before such a score is folded into a commercial product.
What this changes
For the research community, SIMPLE-HF documents a reproducible "distillation" method: use a complex model (here a Transformer) purely as a tool for discovering prognostic variables, then transfer that knowledge into a simple, interpretable score. This approach could apply to other chronic conditions facing the same tension between accuracy and deployability. The next step, flagged by the authors themselves, is external validation in other healthcare systems.
For clinicians, nothing changes today: SIMPLE-HF is a research tool, not prospectively validated or approved for clinical use. But the principle it illustrates is concrete for primary care: a reliable mortality score that needs neither NYHA class nor LVEF could eventually help triage patients at a point where access to a recent echocardiogram is limited — a common situation in primary care.
For patients and the general public, the study illustrates a broader trend: AI in health is not only about building ever more complex models, but also about extracting simpler, more widely usable tools from them. That promise, however, remains conditional on validation in other countries and healthcare systems before any clinical adoption.
Further reading
The paper (published version): Development and validation of a parsimonious AI-based mortality risk score for heart failure, Nouman Ahmed, Nathalie Conrad, Malgorzata Wamil, Ben Omega Petrazzini, Zhengxian Fan, Guyu Zeng, Jie Lian, Shishir Rao, Kazem Rahimi (University of Oxford), npj Digital Medicine, September 26, 2026, DOI 10.1038/s41746-026-03258-1. Preprint with full methodological detail: medRxiv, September 30, 2025.
On other risk scores built and validated on the CPRD database, see our decryption of hospital readmission prediction with a self-supervised autoencoder. On multi-site calibration of a cardiovascular risk score, see our decryption of 10-year ischemic stroke risk.