Chagas disease on the ECG: what an imaging-supervised pre-training actually transfers, and how far
A team from Amsterdam UMC and ETH Zurich aligned an electrocardiogram encoder to a representation space derived from cardiac MRI, using 63,193 paired ECG–CMR examinations from the UK Biobank, then discarded the imaging branch and applied the frozen encoder to Chagas disease detection across roughly 345,000 Brazilian ECGs. Without a single Chagas patient being seen during alignment, the linear probe rises from 0.827 to 0.851 AUROC and from 37.7% to 42.7% sensitivity in the top 5% of predicted risk — ahead of a model trained specifically on Chagas data. But on ELSA-Brasil, the blind cohort furthest from the training data, their AUROC drops to 0.556, the lowest of the four methods compared, on precisely the cohort where the paper claims its best generalization.
The context
Chagas disease is caused by the parasite Trypanosoma cruzi, transmitted mainly by triatomine bugs in Latin America. The World Health Organization classifies it as a neglected tropical disease and estimates between six and seven million people are infected. Its feared complication is cardiac: a Chagas cardiomyopathy that develops years after infection and is, on that continent, one of the leading causes of non-ischemic cardiomyopathy. Its nature is fundamentally structural — myocardial fibrosis, ventricular remodeling, progressive dysfunction.
The diagnostic paradox fits in one sentence. The modality that directly visualizes these abnormalities is cine cardiac magnetic resonance (CMR), regarded as the reference standard for characterizing cardiac structure and function. Yet the authors note, citing published figures, that Latin America has roughly four times fewer MRI scanners per capita than Europe, South-East Asia fifteen times fewer, and Africa up to thirty times fewer. The 12-lead ECG, by contrast, costs almost nothing and exists everywhere — but it records the heart's electrical activity and reflects its structure only indirectly.
Hence the idea, explored over the past two or three years by several groups (MMCL, PTACL, ECCL), of multimodal contrastive learning: take paired ECG–CMR examinations from the same patient, force the model to produce an ECG representation close to that of the corresponding image and far from those of other patients, then discard the imaging. The resulting ECG encoder carries, in principle, a trace of what the imaging would have shown. All these studies, however, were evaluated on the populations and pathologies of their own training cohorts, almost always European or North American. The open question — the one this preprint attacks — is whether these structural representations survive a distribution shift, and more radically whether they retain diagnostic value for a disease entirely absent from the pre-training data.
The method
Two foundation models, only one trained. A foundation model is a large model pre-trained on unlabeled data and intended for reuse across many tasks. On the ECG side, the authors start from ECG-FM, a hybrid CNN–Transformer architecture pre-trained on 1.4 million 12-lead ECGs. On the imaging side, they use CineMA, a multi-view masked autoencoder trained on cine CMR sequences, which maps a short-axis slice and the long-axis views (2-, 3- and 4-chamber) into a shared representation. During alignment the CMR encoder stays strictly frozen: it serves as a fixed target, not as a learning partner.
The diagnosis that changes everything: the raw imaging space is degenerate. This is the paper's most useful observation and it is rarely made elsewhere. CineMA representations at end-diastole (filled heart) and end-systole (contracted heart) have an average cosine similarity above 0.99 — in other words, they are nearly indistinguishable. Static anatomy crushes all the variance and function, that is, motion, disappears. Aligning to these raw vectors therefore amounts to teaching the ECG the size of the heart, not its behavior.
A clinically grounded target rather than a raw one. The authors extract both representations (end-diastole and end-systole), concatenate them, and pass them through a multilayer perceptron supervised to regress seven phenotypes: left ventricular ejection fraction (LVEF), LV end-diastolic volume, LV mass, right ventricular end-diastolic volume, LV global longitudinal strain, atrial fibrillation and myocardial infarction. This perceptron is then frozen. The result is a 256-dimensional bottleneck that concentrates representational capacity on the axes of CMR space most relevant to cardiovascular pathology, stripping away redundant anatomical noise.
The alignment is asymmetric. The loss is InfoNCE — the standard contrastive objective, which rewards bringing matched pairs together and pushing all other pairs in the batch apart, with a temperature parameter (here τ = 0.1) tuning how hard that discrimination is. Here, gradients flow only through the ECG side: the geometry of the imaging space does not move at all, and it is the ECG that must come into line with it. The projection is taken not at the output but at layer 9 of the ECG-FM encoder, a layer chosen by linear probing every intermediate layer against several cardiac phenotypes — thus independently of any Chagas label.
The data. Pre-training on the UK Biobank: 63,193 paired 12-lead ECG and multi-view cine CMR examinations, of which 47,683 for alignment and 15,510 for validation and held-out evaluation, stratified by LVEF, LV mass, sex and atrial fibrillation diagnosis. The ECGs are ten seconds at 500 Hz and are split into two five-second segments, one or the other sampled stochastically as augmentation. Downstream evaluation on the training data released for the PhysioNet/CinC Challenge 2025: CODE-15% and SaMi-Trop combined, roughly 345,000 ECGs with 8,192 Chagas-positive cases — a 2.4% prevalence. For the official challenge protocol comparison, PTB-XL is added as an extra source of negatives (21,801 ECGs). The aligned encoder is used frozen, with a simple linear head trained in five-fold cross-validation, positive-class reweighting and checkpoint selection on AUPRC.
The results
Alignment does add something, and it is measured cleanly. Against exactly the same encoder without alignment — frozen ECG-FM, same probe, same folds — the aligned model rises from 0.827 to 0.851 AUROC (area under the ROC curve: the probability that a randomly drawn positive patient receives a higher score than a randomly drawn negative). More importantly, the challenge's operational metric, Top5%-TPR — the fraction of true positives captured if one keeps only the 5% highest-risk ECGs — rises from 0.377 ± 0.014 to 0.427 ± 0.022, five percentage points.
And this beats a model trained on the disease itself. Jidling et al., who trained an ensemble of fifteen 1D ResNets end-to-end on the full CODE dataset plus SaMi-Trop, reported an AUROC of 0.800. The aligned encoder, which has never encountered a Chagas patient or a Chagas-related image, does better with a simple linear head on top of frozen weights. Under the full challenge protocol (with PTB-XL), the linear probe reaches 0.431 ± 0.005 Top5%-TPR against 0.381 ± 0.003 for the challenge winner's frozen-encoder ablation — the same five-point gap.
The clinical translation brings things back to earth. At 2.4% prevalence, in 1,000 people screened there are about 24 cases. Sending the top 5% of predicted risk for confirmatory serology means testing 50 people. The aligned model catches about 10 of them, the unaligned model about 9: the real gain is roughly one additional case detected per 1,000 people screened. In both cases about forty of the 50 serologies come back negative, and around fourteen cases remain missed below the threshold. This is not a failure — where serological capacity is the limiting factor, correctly prioritizing ten cases across fifty tests has real value — but it is far removed from what "0.851 AUROC" suggests.
On the blind test sets, the picture gets more complicated. The official submission scores 0.269 overall, just behind the top three (0.280 to 0.323), in fourth place. Two results are highlighted by the authors: the highest AUROC on SaMi-Trop-3 (0.773 versus 0.767 for the winner) and the best challenge score on ELSA-Brasil (0.132), presented as the hardest cohort owing to demographic and acquisition differences. What the discussion does not stress: on that same ELSA-Brasil cohort their AUROC is 0.556, against 0.566, 0.567 and 0.626 for the three teams ranked above them. On REDS-II: 0.350 / 0.739.
What is good
The diagnosis of target-space degeneracy, and what they do with it. Observing that a latent space is nearly one-dimensional — end-diastole/end-systole cosine above 0.99 — then refusing to align to it and instead building a 256-dimensional bottleneck supervised on seven clinical phenotypes is a specific, motivated engineering decision, directly reusable by anyone working on ECG–imaging alignment. Most multimodal contrastive papers simply align to whatever the other encoder emits. The asymmetry choice runs in the same direction: freezing the target prevents the contrastive update from distorting the very structural manifold one is trying to transfer.
The control is the right one, and the layer choice is leakage-protected. The five-point gap is not compared to a different model, on other data, with another protocol. It is compared to the same ECG-FM encoder, also frozen, with the same linear head and the same folds. All that changes is the alignment. And layer 9 was selected by probing against generic cardiac phenotypes, explicitly without looking at Chagas labels — closing the door on the most common selection bias in this kind of work, where one picks the layer that performs best on the final task and then presents the result as zero-shot.
The external validation is blind, which is rare. Going through a PhysioNet challenge means accepting evaluation on test sets one has never seen a line of, by a third-party organizer, with a metric fixed in advance. In a subfield where "external validation" most often means testing on a second public dataset one downloaded oneself, the difference is qualitative. The authors further note, in the caption of the official results table, that ranked teams had access to the REDS-II validation set for threshold calibration and that they did not — an inconvenient fact for their standing, which they publish anyway.
What is less good
Population bias cuts exactly against the intended use. The UK Biobank is a cohort of middle-aged to older British volunteers, overwhelmingly white, with a well-documented selection bias toward participants healthier than the general population. The structural prior learned is therefore that of a population where cardiomyopathy is essentially ischemic or hypertensive, where T. cruzi is absent, and whose age pyramid bears no resemblance to a screening population in an endemic region. The authors turn this objection into an argument — "structural features learned from imaging transfer most reliably to populations and settings furthest from the pre-training data" — but this is precisely where the paper is weakest, and no prospective evaluation in Latin America is conducted. The discussion concedes this in one line, deferring it to future work.
The shortcut-learning risk is never audited. The Chagas training pool is an assembly: SaMi-Trop is a cohort of established Chagas cardiomyopathy, hence entirely positive; CODE-15% is a Brazilian telehealth sample, overwhelmingly negative; PTB-XL is German and entirely negative. A classifier can therefore score well by recognizing which cohort a recording comes from — device, sampling, electrode placement, age, filtering — rather than the disease. The paper runs no origin-signature audit: no domain classifier, no negative-control view, no saliency analysis, no per-cohort decomposition of false positives. Yet the collapse from 0.851 in cross-validation to 0.556 on ELSA-Brasil is exactly what this failure mode predicts. We described the same mechanism in dataset fusion in screening mammography.
The most heavily marketed conclusion rests on the metric that moves with the threshold. The claim that the method transfers best to populations furthest from pre-training rests entirely on the ELSA-Brasil challenge score (0.132, best of four), while the AUROC on that same cohort is 0.556, the lowest of the four — barely above chance. The challenge score depends on an operating point; AUROC does not. Since the authors themselves write that they had no access to REDS-II for threshold calibration, the reading "our threshold happened to land well on ELSA-Brasil" is at least as plausible as "our representation generalizes better". One cohort, one submission, no confidence interval: the conclusion is over-read. Three further reservations. The five-point gap comes with no paired statistical test and Table 1 gives no confidence interval on AUROC. No code repository or released weights are announced in the preprint, and both the UK Biobank and the challenge data are access-controlled — exact reproduction is therefore out of reach for an ordinary reader. Finally this is a non-peer-reviewed preprint with no funding or conflict-of-interest statement, and its "Impact in RCS" section suggests a workshop submission on resource-constrained settings rather than a journal article.
What it changes
For the research community. The transferable result is not the 0.851 figure; it is the demonstration that a structural prior learned by alignment to imaging retains value for a pathology absent from pre-training. That shifts the question. Until now, ECG–CMR alignment was sold as a way to improve prediction of phenotypes we already know how to measure. Here it is sold as a way to distil the expensive modality into the available one, once and for all, in a data-rich country, and then ship it elsewhere. If that mechanism holds, it applies to any structural heart disease under-diagnosed for lack of imaging. The first verification to run, however, is the one this paper does not: isolate which phenotypes of the target space actually carry the Chagas signal, and check that it is not the cohort of origin being learned. The question of what medical foundation models really encode connects to what we covered in representational convergence across medical foundation models.
For clinicians. Nothing changes today. A challenge score of 0.269, an AUROC of 0.556 on the most realistic cohort, no prospective evaluation, no regulatory clearance: this is methodological proof of concept. What deserves keeping is the framing, which is intelligent: the point is not to diagnose Chagas from the ECG — diagnosis remains serological — but to prioritize who gets sent for serology when testing capacity is the bottleneck. That is a triage problem, not a diagnostic one, and it is the right way to pose the question. The right comparator, absent here, would be a simple clinical rule: age, region of residence, family history, classical ECG abnormalities of Chagas cardiomyopathy. Until that comparator is beaten on real data, the net gain remains unknown.
For patients and the public. The underlying appeal of this work is a form of technical equity: it proposes learning where imaging is abundant, then deploying where only the ECG exists, without ever requiring MRI at the point of use. That is a more honest direction than asking resource-constrained countries to first acquire rich countries' equipment. But the limit must be seen, and it is structural: the distilled knowledge comes from an older British population, and the one test run on a genuinely different Brazilian cohort yields a ranking barely better than chance. A model shaped by hearts elsewhere is not automatically a model for hearts here. The same tension runs through other AI work on neglected diseases, such as ultrasound screening for periportal fibrosis in schistosomiasis.
Further reading
The preprint: Leveraging Cardiac Imaging to Improve ECG-Based Detection of Chagas Disease in Resource-Constrained Settings, arXiv:2609.08582 [cs.LG], DOI 10.48550/arXiv.2609.08582, submitted 8 September 2026, under CC BY 4.0. Authors: Laura Alvarez-Florez and Daniel Uyterlinde (equal contribution), Samuel Ruipérez-Campillo, Lukas P. A. Arts, Folkert W. Asselbergs and Fleur V. Y. Tjong — departments of Biomedical Engineering and Physics and of Clinical and Experimental Cardiology at Amsterdam University Medical Center, the Quantitative Healthcare Analysis group at the University of Amsterdam, and the Department of Computer Science at ETH Zurich. No funding or conflict-of-interest statement appears in the manuscript. The evaluation framework is that of the George B. Moody PhysioNet Challenge 2025 on ECG-based Chagas disease detection. The main comparator is Jidling et al., PLoS Neglected Tropical Diseases 2023, 17(7): e0011118. The datasets used are the UK Biobank (Sudlow et al., PLOS Medicine, 2015), CODE-15% (Ribeiro et al., Zenodo, 2021), the SaMi-Trop cohort (Cardoso et al., BMJ Open, 2016, 6(5): e011181) and PTB-XL (Wagner et al., PhysioNet, 2022); the reused foundation models are ECG-FM (McKeen et al., JAMIA Open, 2025) and CineMA (Fu et al., arXiv preprint, 2025). On epidemiology and management, the baseline reference is the WHO fact sheet on Chagas disease.