Alzheimer's on MRI: fusing cognitive scores buys 28 accuracy points and degrades what the model learns from the image
Two researchers at the National University of Science and Technology POLITEHNICA Bucharest trained an Alzheimer's classifier on 1,075 T1-weighted MRI scans from the ADNI-1 cohort, then added clinical variables through contrastive learning. Fusing cognitive scores — MMSE, CDR-SB, ADAS11, RAVLT — raises three-way accuracy from 58.7% to 87.3%, but those are precisely the instruments used to assign the diagnostic label in ADNI: the model is not reading the image better, it is recovering the labelling rule. The genuinely useful result lies elsewhere: with architecture, subject split and training held fixed, an encoder aligned to anatomical volumes produces an image representation reaching 73.8% on MCI versus controls, against 52.4% — chance — for the encoder aligned to cognitive scores.
The context
Alzheimer's disease shows up on structural T1 MRI — the sequence that displays anatomy rather than function — as atrophy: brain tissue thins, first in the medial temporal lobe, where the hippocampus and entorhinal cortex sit. For a decade the standard task has been to sort a scan into three categories: cognitively normal (CN), mild cognitive impairment (MCI, the prodromal stage), and established Alzheimer's (AD). Separating AD from CN is comparatively easy, because late-stage atrophy is visible. Separating MCI from CN is hard, and that is the clinically decisive boundary, since it is the window where intervention might still help.
Two problems recur in this literature, and the paper takes on both. The first concerns what the model looks at: a convolutional network fed a whole slice has no reason to restrict itself to brain tissue, and can exploit contrast at the skull, the orbits or the ventricle boundaries. This is textbook shortcut learning — the model picks up a spurious correlation rather than the pathology — the same mechanism documented in the dataset-origin signature in screening mammography.
The second concerns what the model is given. Large cohorts such as ADNI ship rich clinical tables, and recent image–tabular fusion work reports striking gains when those are added to MRI. Except that several columns in those tables — the MMSE, the Clinical Dementia Rating — are among the criteria used to assign the ADNI diagnostic label. Predicting the label from those variables largely amounts to re-deriving the rule that produced it. This is label leakage, a close cousin of the circular leakage already seen with benchmark labels produced by an LLM.
The method
The data. 1,075 T1-weighted MRI scans acquired at 1.5 tesla in the ADNI-1 Screening collection, from 982 subjects — a few participants had more than one screening acquisition. The breakdown is roughly 289 CN, 507 MCI, 279 AD. Every volume was processed with FastSurfer, the deep-learning reimplementation of FreeSurfer, which conforms each scan to a 256³ grid of 1 mm isotropic voxels and produces a DKT cortical parcellation with subcortical labels.
The split. The authors split at the subject level: no patient appears in two sets. This is the point on which Wen and colleagues showed, in 2020, that the Alzheimer's literature had been inflating its numbers by splitting at the slice rather than the patient level. The test set is balanced, 21 scans per class, 63 scans total; binary tasks use 42. The remainder is about 862 training and 150 validation scans.
The image encoder. Deliberately lightweight: an ImageNet-pretrained ResNet18, of which only the last residual block is fine-tuned, applied slice by slice (90 axial slices, indices 62 to 152, resized to 224 × 224). The slice descriptors pass through a single Transformer layer with 8 attention heads — a mechanism that lets each slice weigh information from the others — then mean pooling gives a subject descriptor, projected to a 128-dimensional vector (the embedding, the compact numerical representation of the scan).
The tabular branch and the leakage spectrum. Clinical variables come from the consolidated ADNIMERGE table (16,421 subject–visit rows), organised into five groups ordered by their proximity to the label: cog (CDR-SB, ADAS11, MMSE, RAVLT-immediate — high leakage risk, these are the diagnostic instruments); bio (CSF Aβ, tau, p-tau); pet (FDG, AV45); vol (hippocampus, whole brain, entorhinal, middle temporal — low risk, these are atrophy correlates, not criteria); demo (APOE4, age). Only cog and vol are used: the bio group is populated for only about 14% of rows, which would have required heavy imputation. In cog and vol, fewer than 3% of cells are missing and are replaced by the training-set mean.
The contrastive alignment. The two branches are pulled together by a symmetric CLIP loss — the training objective that pushes a subject's image embedding to resemble their tabular embedding and to differ from those of other subjects in the batch. One honest and rarely reported methodological detail: the loss had to be computed in single precision, because under mixed precision it stayed stuck at its chance value. Two heads are trained on top: an image-only head, a linear classifier that sees no clinical variable at inference, and a fusion head on the concatenation of both embeddings.
The results
Where the classifier looks. YOLOv8 models trained on FastSurfer-derived labels localise the relevant structures with mAP₅₀ of at least 0.95 — up to 0.969 for the medium variant. Compared against those detections, the Grad-CAM maps of the image classifier show two regimes: on many slices it does attend to the ventricles and then the temporal lobes; on a substantial number of others, attention falls on the skull, the neck, the orbits, or outside the head altogether. The authors state that this analysis is qualitative and that they did not quantify the fraction of Grad-CAM mass falling inside the brain.
The fusion that wins is the one that leaks. On the 63 test scans, the three-way task gives 58.7% for image only (95% Wilson interval: 46.4 to 70.0), 87.3% with cognitive-score fusion (76.9 to 93.4) and 73.0% with volume fusion (61.0 to 82.4). The authors explicitly treat 87.3% as a leakage-driven upper bound, not an imaging result.
The main finding: the contrastive target shapes the encoder. This is a controlled comparison. Two training runs identical in architecture, data, subject split, two-phase schedule, hyperparameters and test set. The only difference: the table the contrastive loss aligns the image to. On MCI versus CN, the image-only head — which receives no clinical data at test time — reaches 73.8% (31/42) when the encoder was aligned to volumes, and 52.4% (22/42) when it was aligned to cognitive scores. Twenty-one points apart, with no tabular variable present at evaluation. And the ordering flips under fusion: the configuration that looks best end to end is the one that learned the weaker visual representation.
Cropping to the medial temporal lobe helps. Restricting the input to a 96 × 48 × 64 voxel box centred, subject by subject, on the centroid of eight segmented structures (hippocampus, amygdala, entorhinal and parahippocampal cortex, bilaterally), image-only three-way accuracy rises from 58.7% to 65.1% (41/63, CI 52.8 to 75.7), and MCI versus CN from 73.8% to 81.0%. The gain is larger on MCI than on AD versus CN, consistent with an early, localised signal.
Clinical translation — and its limits. Out of 1,000 people drawn from a population artificially balanced across the three categories, the best image-only model would classify roughly 651 correctly and 349 incorrectly; a coin toss would get 333 right. But the number should not be over-read: on 63 scans the confidence interval spans roughly ±12 points, and the improvement from 58.7 to 65.1% corresponds to four scans. The 21-point contrastive-target gap corresponds to nine scans, with an unpaired two-proportion z-test at p ≈ 0.04 on a single run — the authors themselves call it suggestive, not established.
What is good
The leakage spectrum is explicit, ordered, and built into the protocol. Rather than fusing everything lying around in ADNIMERGE and publishing the highest number, the authors rank variable groups by their proximity to the labelling rule, train one run per group, and declare in advance that the cognitive group is a control upper bound. That is methodological discipline which costs apparent performance and ought to be standard.
The image-only head is reported alongside the fused head. This is the move that makes the paper useful. Without it, the conclusion would have been "cognitive fusion gives 87.3%, ship it"; with it, one sees that this configuration produced the weakest encoder. The resulting recommendation — always report the unimodal head under each target — transfers directly to any multimodal fusion work, and complements what was already known about clinical models where text dominates the image.
The self-criticism is quantified and placed before the conclusion. The authors list the protocol differences that prevent comparison with Huang (63 subject-level scans against 882 slices, ADNI-1 only against a broader selection, one variable group at a time against all groups fused including one containing the baseline diagnosis, single runs against five averaged splits). They write explicitly that the proximity of the numbers shows their reimplementation works, not that it does better. They also report that a 3D encoder tested in a preliminary comparison was about twenty points ahead of their 2D one.
What is less good
The test set is too small for the conclusion it carries — a misleading metric by sample size. Sixty-three scans, twenty-one per class, one run per configuration. With Wilson intervals of ±12 points, nearly every comparison in the paper overlaps. The headline result, the 21-point gap between contrastive targets, rests on nine scans and was not tested pairwise nor repeated over several random seeds — the authors say so, which is honest, but it does not change the evidential status.
No external validation, one cohort, one field strength — population bias. Everything comes from ADNI-1 at 1.5 tesla. ADNI is a research cohort, recruited on a volunteer basis across a limited number of North American sites; nothing indicates that an encoder tuned on its acquisition conventions would hold on a routine 3 T scanner, on an unselected population, or in another health system. The question is the same as for pathology foundation models tested under perturbation: an internal number says nothing about transportability.
The MTL crop shifts the dependency rather than removing it, and some results are missing. The anatomical box requires a FastSurfer segmentation at inference: any segmentation error propagates directly into the crop position, and the pipeline becomes dependent on a computationally expensive external tool. Moreover, the authors write that they trained axial and sagittal variants of the crop but omit their numbers "pending final log extraction" — an omission worth flagging in a paper whose central argument is about not cherry-picking favourable results. Finally, no code or weights repository is mentioned.
What it changes
For the research community, the paper supplies a reproducible and cheap protocol: rank auxiliary variables along a leakage axis, train one run per group, and systematically publish unimodal performance alongside fused performance. The observation that a leaky contrastive target degrades the image encoder — not merely that it inflates the final score — is new and deserves testing elsewhere: if it holds, part of the gains reported in the image–tabular fusion literature measures how easy the table is, not how good the image model is. The decisive test is simple: take existing benchmarks and publish the unimodal head.
For clinicians, nothing changes today. A 65% three-way accuracy on 63 scans from a research cohort is not a tool, and the authors nowhere claim otherwise. What does transfer is a reading grid: when a publication announces that a "multimodal" model reaches 87% on Alzheimer's diagnosis, the first question is whether the MMSE or the CDR are among the inputs. If so, the number mostly measures the internal consistency of the labelling protocol.
For patients and the public, one counter-intuitive point is worth keeping. An AI that correctly says "this person has Alzheimer's" nine times out of ten has not necessarily read the brain: it may have read the cognitive test result that was used to make the diagnosis, which diagnoses nothing new. The value of an imaging model is measured by what it adds to the clinical examination, never by its ability to reproduce it.
Further reading
The paper: Anatomical Grounding and Leakage-Aware Multimodal Contrastive Learning for Alzheimer's Disease Classification from Structural MRI, Paul-Gabriel Nicolae and Irina Mocanu, Computer Science and Engineering Department, National University of Science and Technology POLITEHNICA Bucharest, posted to arXiv on 14 September 2026 (arXiv:2609.15888, cs.CV), 13 pages, DOI 10.48550/arXiv.2609.15888. Preprint, not peer-reviewed as of this decryption's publication date.
Data: the ADNI cohort (Alzheimer's Disease Neuroimaging Initiative), ADNI-1 Screening 1.5 T collection and the consolidated ADNIMERGE table. ADNI data collection and sharing are funded by the National Institutes of Health (grant U01 AG024904) and the Department of Defense (W81XWH-12-2-0012). Experiments ran on the FEP cluster of POLITEHNICA Bucharest, with a twelve-hour per-job limit that the authors cite as the constraint preventing longitudinal multi-phase training.
Tools: FastSurfer (Henschel et al., NeuroImage, 2020) for segmentation and DKT parcellation; YOLOv8 by Ultralytics, AGPL-3.0 licensed, for detection and instance segmentation; Grad-CAM (Selvaraju et al., ICCV 2017) for attention maps; CLIP (Radford et al., ICML 2021) for the contrastive objective; TabTransformer (Huang et al., 2020) for the tabular branch. No code or weights repository is indicated in the paper.
Background references cited by the authors: Wen et al., Medical Image Analysis, 2020, on data leakage in CNN-based Alzheimer's classification; Petersen et al., Neurology, 2010, for the ADNI diagnostic criteria; Huang, ICCVW 2023, for the contrastive fusion work whose protocol this paper adopts and corrects; Hager et al., CVPR 2023, for image–tabular contrastive learning.