Causal multi-modal AI predicts who benefits from chemotherapy — and who will not
AI-enabled perception of how treatment choice shapes patient trajectories. Causal multi-modal AI turns routine pathology and clinical data into personalized estimates of chemosensitivity, separating risk of recurrence from treatment benefit to help direct chemotherapy to the patients most likely to gain from it.
Abstract
If you were diagnosed with the most common form of breast cancer tomorrow, the first treatment decision your oncologist would make would be nearly automatic: prescribe endocrine therapy for multiple years. The next decision would not be: should chemotherapy be added? For some patients, the addition of chemotherapy to endocrine therapy decreases the risk of the cancer recurring. Many others will not derive benefit from chemotherapy, and they spend months on an ineffective drug while suffering from toxic health and financial effects. In today’s clinic, your oncologist would order a genomic test, read the risk score, and make the decision whether to add chemotherapy. What that score does was never designed to tell you is whether you would respond.
Today we are releasing three preprints describing a causal multi-modal AI model, called CTX, that generates personalized chemosensitivity predictions. From a routine pathology slide and standard clinical variables, CTX predicts the gain in five-year recurrence-free probability from adding chemotherapy for a specific patient. CTX was developed by chaining a pathology foundation model together with a causal inference model, using data from 9,141 patients from twelve cohorts from nine countries. Here, we describe this new methodology. We report model performance in held-out data from geographically diverse real-world cohorts (1,994 patients from five cohorts in three countries) and in two phase 3 clinical trials (TAILORx, UNIRAD). that were held out from training. Lastly we evaluate the potential of AI-enabled therapeutic decision support, biological correlates of model predictions, and generalizability with zero-shot transfer to non-breast cancer types.
The clinical problem
HR+/HER2− disease comprises ~70% of the 2.4 million annual breast cancer diagnosed globally. The genomic assays that guide the decision to add chemotherapy including the 21-gene, 50-gene and 70-gene assays, are prognostic tools. These tools were developed to estimate the risk of recurrence, and were only later applied to indicate chemotherapy decision-making. But recurrence risk and chemosensitivity are distinct properties of a tumor.
Let’s illustrate this by thinking about what “chemotherapy benefit’ means for an individual with breast cancer (Figure 1). The green curve is her risk of recurrence on endocrine therapy alone. The gold curve is her risk if chemotherapy is added. The gap between the two curves, therefore, is her personalized chemotherapy benefit.
A prognostic score cannot tell apart two high-risk patients whose tumors differ in how they respond to chemotherapy — one of whom would benefit, and one of whom would not (Figure 2). A low-risk, low-benefit patient may be spared chemotherapy. A high-risk, high-benefit patient should receive it. A high-risk patient who gains little from standard chemotherapy is the one who may need something else entirely – escalation to an additional agent. We reasoned that a causal multi-modal AI model could learn to predict these three elements separately using routine clinicopathological information.
A causal AI solution
Ideally, a biomarker intended to guide therapeutic decision-making should satisfy four properties:
- Developed to estimate the benefit of treatment. A risk score fails this by construction, even if trained with infinite data. It simply doesn’t predict the correct quantity.
- Validated by passing the interaction test. The effect of treatment must depend on the biomarker's value in a survival model.
- Scalable through running on what is already collected. Needing more tissue and genomic testing limits who can get tested.
- Explainable by grounding in known markers. It can be interpreted by the oncologist acting on the result and by the patient.
Current risk-based genomic assays partially satisfy the second and fourth criteria. CTX was built to satisfy all four.
Estimating what would happen under endocrine therapy versus chemoendocrine therapy is a causal inference problem. Advanced machine learning methods approach this problem in observational datasets by learning a representation in which patients who received either treatment look alike, then predict outcomes separately under each treatment. However, these causal methods have not yet been successfully applied in oncology, and we think the reason is that the inputs are insufficiently expressive and the scale of data used in the prior efforts was inadequate.
Neural networks yield expressive representations when trained using the standard supervised learning paradigm. However, they require large labelled datasets, which are in short supply for oncology, as the images are too large (gigapixels) for human experts to annotate. Foundation models overcome this limitation through self-supervised learning. These methods enable the learning of expressive representations from massive unlabeled datasets. We developed our own foundation model, called Falcon, using >2 billion patches taken from pathology slides (Figure 3).
The causal module of CTX receives a compact description of each slide from Falcon and routinely collected clinical variables from the patient’s chart, such as age and tumor size. The causal module entails two predictors: one asks what would happen under endocrine therapy alone, the other what would happen if chemotherapy were added. During training a penalty is applied whenever the model’s internal picture of the two treatment groups drifts apart, so it does not learn biases present in the real-world data. Ultimately, the trained model generates CTX-τ as the difference between the two predictions at five years, a personalized chemosensitivity prediction (Figure 4).
Robust performance across geographically diverse real-world cohorts
We tested CTX in the held-out evaluation set – five cohorts, 1,994 patients, three countries.
For CTX-prognostic we assessed model calibration – the extent to which predicted event probabilities matched the observed event rates (Figure 5). CTX-prognostic's five-year recurrence probability predictions exhibited near-perfect calibration across the full range of predicted risk (slope = 1.00, 95% CI = [0.82, 1.18]; intercept = 0.00, 95% CI = [-0.01, 0.01]).
Next, we tested whether CTX-τ predicted chemotherapy uses the gold-standard measure – the treatment-by-biomarker interaction test. This was run with inverse propensity weighting to adjust for the fact that patients who received chemotherapy were sicker to begin with. The treatment-by-biomarker interaction for CTX-τ was significant (p = 8.11 × 10⁻⁵, Figure 5), indicating strong utility for predicting chemotherapy benefit. We further examined this by stratifying patients by CTX-τ score into subgroups (Figure 5): among low-benefit patients, chemotherapy had no detectable effect (HR 1.40, 0.77–2.57, p = 0.27), whereas in patients identified by CTX-τ as high-benefit, adding chemotherapy more than halved recurrence (HR 0.42, 0.20–0.88, p = 0.022).
Importantly, we benchmarked CTX-τ head-to-head against established clinical, pathological, and molecular biomarkers to determine which actually predict who benefits from chemotherapy (Figure 6). Standard clinicopathological features – age, stage, grade, histological subtype, and Ki67 – showed no significant treatment-by-biomarker interaction. The Oncotype DX recurrence score did (p = 0.025), consistent with its known association with chemosensitivity. CTX-τ, however, displayed by far the most robust interaction across the full evaluation set of 1,994 patients (p = 2.1 × 10⁻⁶). Because Oncotype DX scores were only available for a subset of patients, we also compared the two directly within those same 983 patients: CTX-τ again outperformed Oncotype DX (p = 9.7 × 10⁻⁴ versus p = 0.025).
Performance in randomized clinical trials, TAILORx and UNIRAD
In real-world cohorts allocation to chemotherapy is non-random, which bakes in the biases of clinical judgement. To definitively test CTX as a predictive biomarker, we locked the model and evaluated it in two randomized clinical trials.
The TAILORx trial was set up to explore chemotherapy de-escalation. The trial randomized patients with low clinical risk (node-negative) and intermediate genomic risk to endocrine therapy alone or chemoendocrine therapy, finding endocrine therapy alone was noninferior to chemotherapy on average for this population. Based on routine clinicopathological information and histopathology slides, CTX-τ predictions accurately identified 11% of patients as high-benefit, who derived a significant survival gain from chemotherapy (6% absolute benefit in disease-free interval at five-years, p = 0.011, Figure 7). By contrast, chemotherapy did not measurably affect survival outcomes in the remainder of patients identified as low-benefit by CTX-τ (0% gain, p = 0.83). Notably, CTX-τ displayed a significant treatment-by-biomarker interaction (p = 0.016) while the genomic score did not (p = 0.20).
The UNIRAD trial asked a different question, exploring escalation of therapy. It randomized patients with high clinical risk (node-positive) to adjuvant endocrine therapy with or without an experimental mTOR-inhibitor drug (everolimus). As nearly all patients received chemotherapy, we analyzed whether patients that CTX-τ identified as low-benefit from chemotherapy may derive benefit from escalation to the experimental drug. Indeed, in the CTX-τ low-benefit group, everolimus significantly improved outcomes (gain of 10% in five-year disease-free survival, p = 0.005, Figure 8). In the high-benefit group, it did not (6% loss, p = 0.48). Again, the treatment-by-biomarker interaction was significant (p = 0.01).
Taken together, the two trials provide prospective data on both ends of the risk spectrum. In TAILORx, CTX-τ found chemotherapy benefit that the trial's own biomarker had missed. In UNIRAD, absence of benefit marked out patients for whom escalation of therapy paid off.
AI-enabled therapeutic decision support
Chemotherapy is overprescribed in HR+/HER2− breast cancer even in modern clinical practice, with the use of genomic assays. An unnecessary course of adjuvant chemotherapy means months of treatment, a meaningful risk of hospitalization for toxicity, and lasting side effects such as neuropathy and cardiac damage. It also carries a financial toll that outlasts the treatment itself.
So we asked a simple question: if chemotherapy had been allocated according to CTX-τ rather than contemporary clinical judgement, what would have happened?
We order patients in the evaluation set by their chemotherapy benefit, as predicted by CTX-τ. We then calculated the recurrence rate that would be expected if chemotherapy had been given to the top-ranked patients. This is compared to the observed outcomes in the same cohort. In our evaluation dataset (n = 1,994 patients, five cohorts), 41% of patients received chemotherapy, and 6% suffered a cancer recurrence at five-years after diagnosis (Figure 9). A CTX-guided strategy reached the same 6% while administering chemotherapy to only 10.5% of patients. That means roughly three in four women who received chemotherapy could have been spared it. AI-enabled therapeutic decision support could substantially reduce overtreatment without compromising outcomes.
Explainable signatures of chemosensitivity
“Black box” AI models are met with distrust and poor adoption due to lack of clinician insight into what is triggering the model's predictions, even when they outperform standard tools.
To address this, we looked at what the model was seeing in the tissue. Each slide is divided into thousands of small image patches, and the pathology foundation model turns each one into a numerical description. We grouped patches with similar descriptions into clusters – each with similar morphology – and ranked clusters by CTX-τ score. We then asked three board-certified breast pathologists to examine patches from each cluster and describe what they saw. They described the high-benefit clusters as nests of invasive carcinoma with high nuclear grade, dense cellularity and visible mitotic figures (Figure 10). These are known features of chemo-sensitive tumors. Moreover, they described the low-benefit clusters as fibrotic stroma with little tumor, which is consistent with chemo-resistance.
We then examined the relationship between the CTX-τ score and genomic information. High-benefit tumors were enriched for signatures of proliferation, DNA-damage response, homologous recombination deficiency and interferon signaling. By contrast, low-benefit tumors exhibited programs of estrogen response, epithelial–mesenchymal transition and angiogenesis.
Taken together, tumors identified as high-benefit by CTX-τ display hallmarks of susceptibility to chemotherapy-induced cytotoxicity, concordant across morphological and molecular information.
Zero shot transfer beyond breast cancer – towards a world model for oncology
CTX was trained only on breast cancer, but has learned to predict chemotherapy treatment effects that go beyond breast. When we applied the frozen model to 6,692 patients across sixteen other cancer types, its risk predictions were associated with survival in eleven, most strongly in endometrial carcinoma, which shares estrogen biology with breast cancer, and other adenocarcinomas (Figure 11).
Crucially, CTX-τ correctly ordered chemotherapy benefit in five of the seven cancers treated with platinum- or taxane-based regimens, meaning it mostly worked for treatments which share mechanisms with breast chemotherapy. CTX-τ mostly failed in cancers treated with alkylators or with targeted and immune therapies. We find the failures informative – a score that measured general tumor aggressiveness would order benefit for alkylators as readily as for taxanes. CTX does not, as it learns a treatment-specific predictive signal.
This zero shot transfer distinguishes our result from prior pan-cancer analyses of AI pathology models, which report performance by re-training within each cancer type. That our breast-trained causal model identifies risk and chemosensitivity in the majority of cancer types is a substantially stronger form of generalization.
In the future, the causal multi-modal AI architecture could be extended to learn out-of-domain effects for other therapies and cancer types with minimal additional data – serving as a causal ``world model'' for oncology.
References
- Biswas D, Berrevoets J, McClean A, Bao L, Park J, Zeng KG, Cappadona J, Tang C, Liu C, Machura B, Wu Y, Speirs V, Soliman H, Bhargava R, Kabraji S, Khoury T, Page D, Piening B, Bifulco C, Meurs C, Westenend P, Chabaud S, Lemonnier J, Cottu PH, Dalenc F, Andre F, Penault-Llorca FM, Bachelot T, Howard F, Esteva FJ, Kalinsky K, Pusztai L, Witowski J, Geras KJ. Causal multi-modal AI for personalized chemosensitivity prediction. arXiv. 2026
- Chan N, Tang C, Biswas D, Park J, Zeng K, Cappadona J, Liu C, McClean A, Berrevoets J, Bao L, Machura B, Gray RJ, Wang V, Witowski J, Geras KJ, Sparano JA; TAILORx Investigators. An AI model identifies chemotherapy benefit in node-negative HR+/HER2− breast cancer patients from TAILORx, a phase 3 randomized clinical trial.
- Bachelot T, Chabaud S, Lemonnier J, Cottu PH, Dalenc F, Howard FM, Pusztai L, Tang C, Biswas D, Zeng K, Witowski J, Geras KJ, André F, Penault-Llorca FM. A causal multi-modal AI model stratifies residual risk and identifies candidates for treatment escalation in node-positive HR+/HER2− early breast cancer.
- Ataraxis AI. Bigger and better: what Falcon shows us about scaling pathology foundation models
Cite this post
Copy BibTeX@misc{ataraxis2026ctx,
title = {Causal multi-modal AI predicts who benefits from
chemotherapy---and who will not},
author = {Biswas, Dhruva and Geras, Krzysztof J.},
year = {2026},
howpublished = {Ataraxis AI},
url = {https://www.ataraxis.ai/causal-ai-chemotherapy-benefit}
}