Predicting who actually benefits from chemotherapy: a causal multi-modal model in breast cancer, without a single randomised patient

Thirty-four authors led by Ataraxis AI, a New York company, present CTX: a model that combines the haematoxylin-and-eosin slide with routine clinical data to estimate, patient by patient, the absolute gain chemotherapy brings in early HR+/HER2− breast cancer — 9,141 patients in development, 1,994 in a held-out evaluation set, ten countries. The score beats Oncotype DX on the reference test for predictive value, and the authors calculate that model-guided allocation would cut the number of chemotherapy courses fourfold at an unchanged recurrence rate. But no patient was ever randomised, the confidence intervals betray a very small number of recurrences, and the figure in the abstract — a 30% reduction — is not the one the results section gives.

The context

Roughly 70% of breast cancers are HR+/HER2−: they express hormone receptors (oestrogen, progesterone) and do not express the HER2 receptor. All these patients receive adjuvant endocrine therapy after surgery, depriving the tumour of the oestrogen signalling it depends on. The open question, for twenty years, has been who should get chemotherapy on top. The landmark trials showed that adding it improves survival on average — but the average conceals wildly unequal situations, and chemotherapy is expensive in toxicity: neutropenia, lasting peripheral neuropathy, anthracycline cardiotoxicity, early menopause, prolonged fatigue.

To decide, guidelines rely on genomic signatures: the 21-gene recurrence score (Oncotype DX), the 70-gene signature (MammaPrint), the 50-gene score (Prosigna/PAM50). All were designed and validated as prognostic tools — they estimate recurrence risk — then repurposed as predictive ones, that is, as indicators of expected treatment benefit. The distinction is far from a vocabulary quibble. A prognostic marker says "this tumour is dangerous"; a predictive marker says "this tumour will respond to that drug". TAILORx, RxPONDER, MINDACT and more recently OPTIMA showed that chemotherapy can be withheld from patients classed as low genomic risk without worsening outcomes on average. They never showed that high genomic risk identifies chemosensitive patients: among the high-risk, some gain nothing from chemotherapy and would be better served by escalation with other agents, such as CDK4/6 inhibitors.

That gap is what the paper attacks. Rather than predicting risk and reading it as benefit, it proposes to estimate directly, for each patient, the difference between two futures: probability of recurrence under endocrine therapy alone, probability of recurrence under endocrine therapy plus chemotherapy. The difference is the personalised benefit. And because this reading is done on a haematoxylin-and-eosin slide already produced in routine care for every breast cancer, it requires no extra molecular assay, no extra tissue, and none of the several thousand euros the commercial signatures cost.

The method

The data. A multi-national observational dataset of 11,135 patients with early breast cancer, drawn from 17 cohorts across ten countries. Split: 9,141 patients in development (12 cohorts, nine countries), 1,994 HR+/HER2− patients in a held-out evaluation set (5 cohorts, three countries), including French cohorts (UNIRAD, Centre Léon Bérard, Institut Curie, Gustave Roussy, Centre Jean Perrin and IUCT-Oncopole appear among the affiliations). Each patient contributes harmonised clinical variables and at least one whole-slide image. Chemotherapy was received by 47.4% of development patients and 40.7% of evaluation patients. The endpoint is five-year recurrence-free interval (RFI): time from diagnosis to first local-regional or distant recurrence, with death treated as censoring.

The pathology encoder. Slides are tiled into 224 × 224 patches encoded by Falcon, Ataraxis AI's in-house foundation model — successor to Kestrel — a giant vision transformer pretrained by self-supervision on over two billion patches using a DINOv2-like recipe: the model is taught to produce the same representation of a tissue region seen under two different transformations, with no labels at all. Falcon stays frozen during CTX development; its outputs are aggregated into a single slide-level representation via multiple instance learning (bag-level learning: you do not know which patch carries the signal, only that the whole slide is tied to an outcome). Clinical variables run in parallel through a tabular residual network, and the two representations are fused by a gated attention layer.

The causal component. This is the heart of the paper. Because chemotherapy was not randomised, treated patients differ systematically from the others: younger, larger tumours, more involved nodes. A naive model trained on that would mostly learn to recognise who was treated, and would fail badly on the counterfactual scenario — what would have happened to the same patient under the other treatment. The authors therefore use a CFRNet (counterfactual regression network): the network learns an internal representation in which the two treatment groups resemble one another, explicitly penalising the gap between their distributions with a measure called maximum mean discrepancy. On top of that balanced representation, two survival heads predict five-year recurrence risk under each treatment. Hence the two outputs: CTX-prognostic, the risk under the treatment actually received, and CTX-τ, the predicted absolute increase in recurrence-free probability from chemotherapy.

A 2% threshold separates high from low benefit. It is justified: patients report that a 1% gain makes chemotherapy worthwhile, an oncologist consensus panel puts the bar at 3–5%, and the authors take the middle ground while insisting the score stays continuous so each clinician-patient pair can set its own cut-off.

Evaluation. Here too treatment is not randomised in the evaluation set. The authors therefore estimate a propensity score — the probability, given clinical covariates, that a patient received chemotherapy — and reweight all analyses by its inverse (IPW). Measured effect: the absolute average standardised mean difference between groups falls from 0.376 to 0.125, below the 0.20 threshold conventionally regarded as negligible imbalance.

The results

Prognosis. On the 1,994 held-out patients, CTX-prognostic achieves a time-dependent AUC of 0.711 [95% CI 0.663–0.756] at five years and 0.765 [0.708–0.791] at ten — AUC being the probability that the model ranks a patient who recurs above one who does not. Honest discrimination, no more, comparable to the best clinical scores. The score remains an independent predictor after adjusting for age, tumour stage and nodal stage (adjusted HR 2.02 [1.44–2.83] at five years; 2.28 [1.73–3.00] at ten). Calibration — the agreement between predicted probability and observed frequency — is the most solid result of the set: slope 1.00 [0.82–1.18], intercept 0.00 [−0.01, 0.01]. A model that says 8% risk does see 8% recurrences.

Predictive value. Among patients classed low-benefit, chemotherapy changes nothing measurable (HR 1.40 [0.77–2.57], p = 0.270). Among those classed high-benefit, it more than halves the risk (HR 0.42 [0.20–0.88], p = 0.022). The reference test is the treatment-by-biomarker interaction in a Cox model: it is significant after IPW weighting (p = 8.11 × 10⁻⁵) and stays so in continuous form (HR per standard deviation 0.67 [0.57–0.79], p = 2.09 × 10⁻⁶), ruling out a threshold artefact. Ablation shows that clinical features alone (p = 4.41 × 10⁻⁴) and pathology alone (p = 7.90 × 10⁻³) each carry signal, and that adding pathology significantly improves fit beyond clinical data (likelihood ratio test, p = 9.79 × 10⁻³). Against a panel of established markers — age, stage, grade, histology, Ki67 — none reaches significance; only Oncotype DX does (p = 0.0247). In the strict head-to-head on the 983 patients with both scores, CTX-τ retains the stronger interaction (p = 9.68 × 10⁻⁴ versus 0.0247).

Therapeutic decision. This is the figure that will make headlines. In the evaluation set, 40.7% of patients received chemotherapy, for a five-year recurrence-free rate of 93.7%. Ranking patients by decreasing CTX-τ, the authors estimate the same rate would be reached by treating only 10.5% of them — a relative reduction of 74.2%. Alternatively, at an identical treatment rate, the recurrence rate would fall from 6.3% to 5.4%, or −14.3% relative.

In clinical terms. Out of 1,000 patients like this cohort, 407 currently receive chemotherapy and 63 will recur within five years. The CTX-guided strategy would treat 105 of them for the same number of recurrences: 302 chemotherapy courses avoided per 1,000 patients. Conversely, treating the same 407 would spare 9 recurrences per 1,000. These figures belong beside a third one the paper reports without highlighting: across the whole evaluation cohort, the average benefit of chemotherapy is 0.30% in five-year recurrence-free probability. Three patients in a thousand. That tiny quantity, and its heterogeneity, is what the entire exercise is about.

Biology and transfer. Three board-certified breast pathologists, blinded to predictions, reviewed the morphological clusters associated with the score. High-benefit tumours show nests of invasive carcinoma, high cellularity, intermediate-to-high nuclear grade, identifiable mitotic figures; low-benefit ones are dominated by fibrotic stroma and adipose tissue, with little to no invasive carcinoma. Molecularly, in TCGA (n = 348), high-benefit tumours are enriched for proliferation pathways (E2F targets, G2M checkpoint, MYC targets, mTORC1), DNA repair and immune activation (interferon α and γ, cGAS-STING); low-benefit ones for epithelial-mesenchymal transition, angiogenesis and early oestrogen response. Finally, the frozen model applied without any recalibration to 6,692 patients across 16 TCGA cancer types keeps a significant prognostic signal in 11 of 16, and in 6 of the 7 types treated with regimens from the same cytotoxic family.

What is good

The question is the right one, and it is asked cleanly. Explicitly separating prognosis from chemosensitivity, producing two distinct outputs rather than one reinterpreted score, and demonstrating it through the correlation between them — CTX-τ correlates only r = 0.389 with the Oncotype DX score, r = 0.345 with tumour size, r = 0.340 with node count — is the theoretical point of the paper, and it is supported. Even the correlation with CTX-prognostic, produced by the same model, tops out at r = 0.720: the two axes do not overlap.

The evaluation uses the right test and reports its assumptions. The treatment-by-biomarker interaction in a Cox model is indeed the oncology reference for establishing that a marker is predictive and not merely prognostic, and the authors report it dichotomised, continuous, with and without propensity weighting, across five held-out cohorts in three countries. Above all they publish the imbalance before and after reweighting (0.376 → 0.125): stating the size of the problem you claim to fix remains rare, and lets the reader judge. Calibration is reported with slope and intercept, not as a qualitative plot. The comparison with Oncotype DX is made in the same patients, not across two different populations.

Explainability is not a heat map. Three board-certified pathologists, blinded to predictions and outcomes, annotate the morphological clusters; the molecular analysis is run separately, and the authors themselves note it is exploratory since TCGA belongs to the development set. The two readings converge on the same story — proliferation, replication stress, homologous recombination deficiency on one side; oestrogen signature and microenvironment-mediated resistance on the other — which is biologically plausible for cytotoxics acting on dividing cells.

What is less good

No patient was randomised, and weighting cannot repair that. The whole construction rests on the assumption of no unmeasured confounding: that harmonised clinical variables suffice to explain why one patient received chemotherapy and another did not. This assumption is untestable, and in breast oncology it is frankly doubtful. The chemotherapy decision depends in practice on performance status, cardiac and haematological comorbidities, desire for pregnancy, the informed patient's preference, centre habits, year of care — none of which appears in a harmonised multi-country record. The propensity score rebalances only what it observes; it is blind to the rest by construction. A residual standardised difference of 0.125 says balance is good on the included variables, not that confounding has gone. This is the biased comparator failure mode in its most stubborn form: what the model calls chemosensitivity could partly be the trace of unrecorded reasons for treating.

The confidence intervals say events are very few. With a five-year recurrence rate of 6.3% in 1,994 patients, we are looking at roughly 125 events for the whole evaluation set — then split across two benefit categories and two treatment arms. The key result, HR 0.42 [0.20–0.88] with p = 0.022, is compatible with a fivefold risk reduction and with a 12% one. The low-benefit group, HR 1.40 [0.77–2.57], excludes neither benefit nor harm. The numbers-at-risk tables under the Kaplan–Meier curves are IPW-weighted and therefore do not correspond to real head counts, making robustness even harder to judge. The contrast between six-decimal p-values and intervals this wide is itself informative: significance holds, precision does not. This is the same problem already seen in other prognostic models validated on very few events.

The morphology described looks like a measure of tumour content. Low-benefit tumours are, by the reviewing pathologists' own account, dominated by fibrotic stroma and fat "with little to no invasive carcinoma". There are two readings. Either the model captured a microenvironment-driven resistance biology — the authors' reading, consistent with the epithelial-mesenchymal transition enrichment. Or it is mostly measuring how much tumour is present on the digitised slide, that is, a sampling and grossing artefact correlated with tumour burden, hence with risk, hence with expected benefit. This is a classic shortcut learning candidate, and the paper reports no control that separates the two: no tumour-content-matched analysis, no test on slides cropped to the invasive area, no comparison across blocks from the same case. The same trap has been demonstrated in breast imaging, where a model learned the signature of its source site rather than the lesion, and the fragility of pathology foundation models under preparation variation is documented.

Three further points call for caution. Abstract and results do not give the same number: the former announces a reduction in chemotherapy-treated patients "by 30%", the latter computes a move from 40.7% to 10.5%, or 74.2% relative. The two cannot describe the same quantity; the reader does not know which is right, and this is precisely the figure that will be quoted. The model is proprietary: Falcon and Kestrel are Ataraxis AI's foundation models, the first and last authors are affiliated with the company, correspondence runs through a corporate address, and nothing in the text announces public weights. This is far from something an academic lab could reproduce. Finally, the "zero-shot" transfer is presented more generously than it is measured: it uses overall survival rather than recurrence, in TCGA, with the high/low benefit split fixed at the 90th percentile, and it fails in 5 of 16 cancer types — which does not forbid seeing a signal, but does not support the idea of a "universal strategy".

What it changes

For the research community, the paper cleanly installs a methodological distinction computational pathology was missing: stop treating a prognostic score as a proxy for treatment benefit, and learn the treatment effect directly with a balanced-representation network. The demonstration that the H&E slide carries chemosensitivity information not recoverable from clinical variables (p = 9.79 × 10⁻³) is the most reproducible result of the set, and the easiest to replicate elsewhere. It also opens an uncomfortable and useful question: if a causal model trained on observational data yields treatment-biomarker interactions more significant than a test validated in randomised trials, should we read better biology or better-exploited residual confounding? The literature on learning treatment policies from retrospective data suggests the question deserves asking before, not after.

For clinicians, nothing changes today, and that should be said plainly. A predictive chemosensitivity score cannot be adopted on the strength of reweighted retrospective analyses, however careful: Oncotype DX itself entered guidelines only after TAILORx and RxPONDER, that is, after prospective trials in which de-escalation was the arm under study. The evidence required here is of the same kind: a trial where chemotherapy allocation is decided by the model in one arm and by standard practice in the other. Until that exists, CTX is a quantified hypothesis, not a decision tool. What is worth keeping, on the other hand, is the order of magnitude the paper recalls: an average benefit of 0.30% at five years across a whole HR+/HER2− population means the vast majority of chemotherapy courses prescribed in this group bring nothing to the women who receive them. The overtreatment problem is real, whatever the solution.

For patients and the general public, the promise is easy to state and hard to keep: read from a slide already produced, with no extra test and no extra cost, whether chemotherapy is worth going through. The gap between that promise and the state of the evidence fits in one sentence: we know the score separates groups with different outcomes in past records; we do not yet know what would happen to women whose treatment was decided by it. The two case studies in the paper — two clinically indistinguishable patients whose fates followed the prediction — are illuminating for understanding the concept, and the authors themselves state that they establish no accuracy. That is an honest caveat; it deserves repeating every time the result is quoted.

Further reading

The paper: Causal multi-modal AI for personalized chemosensitivity prediction, arXiv:2609.13567 (cs.AI), submitted 11 September 2026, DOI 10.48550/arXiv.2609.13567.

On the causal framework used: counterfactual regression networks (CFRNet) by Shalit, Johansson and Sontag, which introduce the balancing penalty between treatment groups. On the trials that currently define chemotherapy de-escalation in HR+/HER2− breast cancer: TAILORx (NCT00310180), RxPONDER (NCT01272037), MINDACT (NCT00433589), OPTIMA (ISRCTN42400492). On the imaging foundations: DINOv2 for the self-supervision method, and the digital pathology foundation models of which we decrypted the GigaPath case.