Biologically-informed AI predicts treatment response from tumor biopsies
Abstract
Molecular portraits of tumors, typically generated using large amounts of tumor tissue and with research-grade technologies, have been shown to predict therapy response. However, for neoadjuvant therapy (treatment given before surgery) response in breast cancer, predictions must be made with clinical-grade assays fast enough to inform treatment decisions on tiny amounts of tissue, which vastly limits the amount of information available about the tumor’s biology. AI offers a way to extract more predictive information from these limited data and has already shown success across multiple precision oncology applications. However, cohorts of patients with both a digitized pre-treatment biopsy slide and a recorded response to neoadjuvant therapy are typically in the dozens to hundreds of patients – too few to develop deep learning models.
Here, we developed a two-stage AI strategy to solve this problem. First we developed MORPHEUS, an AI model that estimates the expression of 14,773 genes from a routine biopsy slide. Next, we developed NEO, which uses this virtual molecular portrait to predict response to neoadjuvant therapy. NEO demonstrated excellent discriminative performance for neoadjuvant therapy response in 1,412 patients across geographically and molecularly diverse real-world cohorts. Further, we evaluated NEO in two clinical trials, showing how the model’s predictions may be used to help inform a specific treatment decision. A deep dive into MORPHEUS revealed exactly which molecular signals AI sees and learns from in pathology images, findings we verify with expert pathologists. Lastly, we stress-tested NEO across the challenges it will encounter in clinical practice, showing that model predictions were stable regardless of the biopsy site and amount of tissue sampled.
These data indicate that biologically-informed compression can be applied to predict neoadjuvant response in breast cancer, and suggest this approach may be applied to solve other data-sparse prediction tasks across precision oncology.
The clinical challenge – and why it is hard for AI to learn
Neoadjuvant therapy refers to systemic treatment given before surgery, such as chemotherapy or targeted therapy. It is becoming the standard of care for breast cancer. The goal is a pathological complete response (pCR): no invasive cancer left in the breast or lymph nodes at surgery. Patients with tumors that achieve pCR after neoadjuvant treatment may be candidates for less extensive surgery and, in ongoing trials, for less chemotherapy. However, neoadjuvant therapy means months of systemic treatment, often including cytotoxic chemotherapy, before definitive surgery.
About two out of three patients achieve pCR. Estimating a patient's probability of pCR before treatment (Figure 1) could help inform bespoke therapeutic decision making for each individual patient. It could identify likely responders who might be candidates for less intensive neoadjuvant treatment, and flag patients at high risk of residual disease for the consideration of escalation strategies. None of the existing predictors of neoadjuvant therapy response have been deemed sufficiently performant to receive FDA-approval or be recommended in clinical guidelines.
This clinical task has been hard to learn as response prediction is data-poor by nature. A patient can be used to train or evaluate a response-prediction model only if they received neoadjuvant therapy, had a pre-treatment biopsy slide that was digitized, and had their response recorded at surgery. Most breast cancer datasets fail at least one of those requirements.
Ataraxis has assembled data from tens of thousands of early breast cancer patients, the basis for our recurrence-risk and chemotherapy-benefit models. Fewer than 10% of these patients have the required information. Even so, we assembled a total of 2,492 patients with the required information from 14 cohorts across seven countries. To our knowledge, this is the largest collection of its kind.
Yet by the standards of modern AI, this is still small. A whole-slide image contains millions of cells and thousands of image patches. A response label is a single yes or no per patient. Learning which of those visual details matter from so few labels, even starting from a powerful pathology foundation model, is objectively difficult.
A two-stage AI solution
To address these challenges, we devised a two-stage solution (Figure 2).
(1) MORPHEUS learns how tissue morphology relates to tumor biology. It was developed on paired H&E slides and measured gene activity from 8,742 patients across 32 cancer types in The Cancer Genome Atlas. Here, every patient provides 14,773 measurements, one for each gene, rather than one outcome label: a far richer training signal than response data can offer. Once trained, MORPHEUS estimates the expression of those 14,773 genes from the H&E slide alone, without requiring an expensive, tissue-consuming genomic assay. We call these estimates the “virtual molecular portrait” of the tumor.
(2) NEO then learns how to predict pCR, using the virtual molecular portrait and routine clinical variables. NEO was developed using paired H&E slides and neoadjuvant response data from 1,080 patients across five cohorts. Because MORPHEUS has already learned how tissue relates to biology, NEO only needs its pCR labels to relate that biology to response, rather than to learn everything from morphology alone.
Robust performance across geographically and molecularly diverse real-world cohorts
We evaluated NEO on 1,412 patients from nine independent cohorts across five countries. These evaluation cohorts were independent from the data which contributed to model development. The cohorts differed in size, subtype mix, response rate, and slide preparation: a realistic test of whether the model generalizes (Figure 3).
We benchmarked NEO against histopathological biomarkers. Tumor-infiltrating lymphocytes (TILs), immune cells in and around the tumor, are the most studied histopathological predictor of response, and there exist several computational methods scoring them. NEO distinguished patients who achieved pCR from those who did not, far more accurately than four computational TIL biomarkers (AUROC 0.79 versus 0.52–0.58; p < 0.0005 for each; Figure 4). Similarly, NEO outperformed Ki-67, a proliferation-based predictor of neoadjuvant response (AUROC 0.87 versus 0.65; p = 0.012; Figure 4).
Performance in HER2-positive de-escalation clinical trials: PHERGain and PHERGain-2
Dual HER2 blockade with trastuzumab and pertuzumab (HP) plus chemotherapy is the standard neoadjuvant treatment for HER2-positive early breast cancer. It achieves high response rates, but not every patient needs chemotherapy. PHERGain and PHERGain-2, two phase II trials run by MEDSIR, tested whether selected patients could reduce or skip chemotherapy without compromising outcomes. Both trials showed that de-escalation is feasible. However, neither could predict which patients were the right candidates for de-escalation before treatment started. We asked whether NEO, applied to the pre-treatment biopsy, could fill that gap.
We generated NEO scores for 689 patients across the two trials. Neither trial was used to develop or train the model. We assessed NEO against pCR, a key early signal that a de-escalated regimen has done its job.
PHERGain enrolled patients with stage I–IIIA disease and randomized them to standard chemotherapy plus HP (TCHP) or to an adaptive strategy. In the adaptive arm, patients started on HP alone and added chemotherapy only if an early PET scan or surgery showed an inadequate response. Across the trial, higher NEO scores meant a higher likelihood of pCR (p < 0.001). Using thresholds pre-specified for HER2-positive disease, NEO sorted patients into groups with varying response rates (Figure 5, left):
- Low-probability: 19% achieved pCR.
- Medium-probability: 41% achieved pCR.
- High-probability: 48% achieved pCR.
High-probability patients had nearly four times the odds of achieving pCR compared with low-probability patients (OR = 3.92, 95% CI 1.86–8.95, p < 0.001), and this pattern held after adjusting for treatment arm and hormone receptor status.
The clearest signal came from patients whose early PET scan showed little response to HP. None of the 14 patients NEO had flagged as low-probability went on to achieve pCR. In the medium- and high-probability groups, 37% and 44% did respectively (Figure 5, left).
PHERGain-2 was a harder test. It enrolled a population already selected for favorable biology: HER2 IHC 3+, node-negative tumors of 5–30 mm. PHERGain-2 was a single-arm trial where all patients received neoadjuvant HP with no chemotherapy. However, even in this already-enriched group, NEO separated outcomes sharply (Figure 5, right):
- Low-probability: 30% achieved pCR.
- Medium-probability: 62% achieved pCR.
- High-probability: 70% achieved pCR.
High-probability patients had more than five times the odds of achieving pCR compared with low-probability patients (OR 5.59, 95% CI 2.64–12.39, p < 0.001). After adjusting for hormone receptor status, the effect grew stronger (OR 6.85, 95% CI 2.91–16.84, p < 0.001).
Taken together, PHERGain and PHERGain-2 point to where NEO is most useful. While the separation between medium and high probability was modest, the low-probability group stood apart in both trials. These patients are candidates for de-escalation based on their clinical profile, yet their tumors rarely achieve pCR on HER2-targeted therapy alone. In PHERGain-2, standard eligibility criteria identified a population expected to do well without chemotherapy. Within that population, NEO placed about one in nine patients in its low-probability group, and only 30% of them achieved pCR. Identifying those patients before treatment is the missing piece for safe chemotherapy-free strategies. NEO does it from a slide that is already collected at diagnosis.
What MORPHEUS sees
Before NEO predicts response, MORPHEUS first predicts the tumor's underlying biology. Because every NEO prediction rests on the virtual molecular portrait, we can ask whether it captures real biology, and whether that biology drives NEO's predictions. Across 23 cancer types, the virtual molecular portraits were most faithful for genes involved in immune activity, stromal remodeling (changes in the connective tissue around the tumor), and proliferation (cell division). These are programs that leave visible traces in tissue.
Pathologist review confirmed that - for human-readable labels - MORPHEUS is seeing the same thing as human experts. Using MORPHEUS to score individual image patches for epithelial–mesenchymal transition (EMT), a process in which tumor cells loosen their attachments to one another and become more mobile, we scored patches from biopsy slides. The patches MORPHEUS scored lowest and highest for EMT were reviewed by two board-certified pathologists, who confirmed that low-predicted-EMT patches were mostly carcinoma with epithelioid cells, and high–predicted-EMT patches were mostly connective tissue made of elongated, spindle-shaped cells (Figure 6).
MORPHEUS was trained only on a single measurement per gene averaged over a whole tissue sample. It was never told where in the tissue a gene is expressed. We asked whether MORPHEUS can nonetheless recover spatial patterns, using two public breast cancer slides in which gene activity had been measured in each location in the tissue. Neither slide was seen during training. We divided each slide into thousands of small patches, and asked MORPHEUS to estimate each gene's expression in every patch separately. Coloring each patch by its estimated activity gives an estimated map of the gene across the slide. To evaluate MORPHEUS, we then compared this map with the gene expression measured in the same patches (Figure 7), using a Spearman correlation: 1 means they rank the patches identically, and 0 means they are unrelated.
For IL7R, a gene active in immune cells, the estimated and measured maps agreed closely (Spearman correlation of 0.65 across 8,068 patches). For TAGLN, a gene active in a type of connective-tissue cells, they also agreed (0.54 across 3,614 patches).
Robust to intratumoral sampling and with minimal biopsy tissue
Breast tumors are biologically heterogeneous - comprising different cell populations (cancer, immune, stromal) and different cancer cell sub-populations. Neoadjuvant therapy decision-making is informed by a needle biopsy, which samples less than one percent of a tumor. A biomarker whose result depends on where the needle landed, or on how much tissue was collected, is ineffective. We stress-tested the robustness of NEO to sampling site and tissue availability.
While NEO's predictions differ widely from patient to patient, they were highly consistent from different slides of the same patient (Figure 8, top). Zooming in to look at four example slides, we scored each biopsy core separately, and the cores gave similar predictions to one another and to the whole slide (Figure 8, bottom).
Biopsies don’t always capture the same amount of tumor. So how much tissue does NEO need to make a consistent prediction?
We tested an image-only version of NEO on 70 large slides, repeatedly sampling random subsets of tissue and comparing the results with predictions from the full slide. With just 500 patches, about 6 mm² of tissue and less than 5% of the full slide, predictions changed by a median of only 3.4% relative to the full-slide prediction. Only at the smallest sample tested, 50 patches or about 0.6 mm², did the change grow larger, to 11% (Figure 9).
These results indicate that NEO produces similar predictions even when much less tissue is available.
Biologically informed compression may generalize to data-sparse applications in precision oncology
NEO shows that a routine diagnostic slide can yield a virtual molecular portrait of the tumor, and that a model built on this portrait accurately predicts response to neoadjuvant therapy, outperforming established tissue biomarkers, and stays stable across biopsy regions and tissue amounts. The PHERGain trials provide high-quality, prospectively collected data, suggesting that NEO could provide utility for the sort of clinical decision oncologists and patients face in practice–chemotherapy de-escalation in HER2+ tumors.
The broader lesson concerns data. Clinical outcome labels will remain scarce for many of the questions oncology most needs to answer. Learning the relationship between morphology and molecular biology once, from large paired datasets, and reusing it for each new clinical endpoint, may be a way to build useful models where labeled patients number in the hundreds rather than the tens of thousands.
Research behind this post
- Jungkyu Park, Dhruva Biswas, Joseph Cappadona, Cerise Tang, Ken G. Zeng, Bartosz Machura, Chuwen Liu, Paolo Tarantino, Coral Omene, Francisco J. Esteva, Rohit Bhargava, Marcin Braun, Kamila Paździerz, Jakub Czerwiński, Hanna Romańska-Knight, Albert Grinshpun, Bareket Daniel, Michele Buchinger, Frederick Howard, Piotr Wysocki, Brie Chun, Freya Schnabel, Rich Caruana, Jan Witowski, Krzysztof J. Geras. Transcriptome-informed multi-modal AI for predicting neoadjuvant therapy response from breast cancer biopsies. arXiv. 2026.
- Manuel Ruiz Borrego, Javier Cortés, Agostina Stradella, Begoña Bermejo, Santiago Escrivá, Cristina Reboredo, Cinta Albacar, Laia Garrigós, Sherko Kümmel, Marco Colleoni, Geraldine Gebhart, Florence Dalenc, Khaldoun Kerrou, Sofia Braga, Peter Schmid, Frederik Marmé, Serena Di Cosimo, Giulia Notini, Elena Martínez-García, Ana Amaya-Garrido, Daniel Alcalá-López, Leonardo Mina, Jose Rodriguez-Morató, Cerise Tang, Dhruva Biswas, Jungkyu Park, Bartosz Machura, Jan Witowski, Krzysztof J. Geras, Mario Mancino, Jose Manuel Pérez-García, Antonio Llombart-Cussac. Multi-modal artificial intelligence predicts pathological response and prognosis in HER2-positive early breast cancer treated with chemotherapy de-escalation strategies: analyses from PHERGain and PHERGain-2 Trials. medRxiv. 2026.
Cite this post
Copy BibTeX@misc{ataraxis2026neo,
title = {Biologically-informed AI predicts treatment response
from tumor biopsies},
author = {Park, Jungkyu and Biswas, Dhruva and Tang, Cerise and
Geras, Krzysztof J.},
year = {2026},
howpublished = {Ataraxis AI},
url = {https://www.ataraxis.ai/biologically-informed-ai-neoadjuvant-response}
}