Neoadjuvant chemotherapy (NACT) is increasingly used for early-stage breast cancer eligible for chemotherapy, offering the advantage of in-vivo assessment of treatment efficacy and allowing surgical de-escalation based on tumor response. The concept of response-guided treatment, where the regimen is adjusted based on early response signals, requires reliable prediction tools that are available from the outset of treatment.
Achieving a pathological complete response (pCR), defined as the absence of residual invasive cancer in the breast and lymph nodes after NACT, is an established surrogate for improved survival. Identifying who will achieve pCR before treatment or early during treatment would allow oncologists to optimize regimens, spare non-responders from ineffective toxicity, and direct likely responders to treatment de-escalation if appropriate.
While MRI and PET/CT have been explored for treatment response prediction, digital mammography (DM) remains the most accessible imaging modality worldwide. DM images encode information about both the tumor and the surrounding breast parenchyma through pixel-level differences in tissue density and architecture. The hypothesis of this study is that AI can decipher these patterns as predictors of treatment response in a way that human visual inspection cannot.
To the authors' knowledge, this was the first study to investigate applying AI to baseline pre-treatment mammograms specifically to predict pCR to NACT in breast cancer patients, making it a proof-of-concept investigation with implications for a widely accessible, low-cost predictive tool.
The study included 453 female breast cancer patients receiving NACT across two cohorts at Swedish hospitals: a retrospective cohort (258 patients from 2005 to 2016) and a prospective cohort (195 patients from 2014 to 2019). The large majority (90 percent) received an epirubicin-cyclophosphamide based regimen followed by taxane, with HER2-positive patients additionally receiving trastuzumab and/or pertuzumab.
Of the 453 patients, 400 were used for model training and validation in an 80/20 split, while 53 patients were reserved as a completely separate held-out test set selected before any model development began. This strict separation between training and test data prevents optimistic bias in performance estimates. A total of 2,514 mammographic images from three different scanner vendors (Siemens, GE, Philips) were analyzed.
The deep learning system operated in two sequential steps. First, a detection transformer (DETR) model was applied to each mammogram to locate and bound the tumor. Second, two image patches, one centered on the detected tumor and one extracted from the equivalent position in the contralateral cancer-free breast as a reference, were fed into a ResNet18-based classification network with parallel pathways that concatenated their features to predict pCR probability.
This dual-patch design forced the model to focus on locally relevant information rather than irrelevant background patterns, and the use of the contralateral breast as a reference provided implicit normalization for individual patient characteristics such as overall breast density and imaging conditions. Histogram equalization was applied to all images to standardize contrast across different scanner vendors.
Of the 453 patients, 95 (21 percent) achieved pCR. The median tumor size was 30 mm, and 68 percent had lymph node involvement at diagnosis. Consistent with established biology, pCR rates varied dramatically by molecular subtype. HER2-positive patients had the highest pCR rate (51 of 134, or 38 percent of cases achieving pCR represented 54 percent of all pCR cases). Luminal A patients achieved pCR in zero cases out of 45, confirming the well-established low chemosensitivity of this subtype.
The pCR group had substantially different molecular characteristics from non-pCR patients: 74 percent were estrogen receptor negative vs. 29 percent among non-pCR, and 83 percent were progesterone receptor negative vs. 41 percent. These biological differences mean the mammographic patterns associated with pCR likely partially reflect the underlying molecular phenotype of tumors more prone to complete responses.
Mammographic density distribution was similar between pCR and non-pCR groups, with approximately half of patients in both groups falling into heterogeneously dense (BI-RADS C) categories. This suggests that overall mammographic density alone does not capture the information relevant to treatment response, consistent with prior work from this group showing inconclusive results when using density alone as a predictor.
On the 53-patient independent test set, the AI model achieved an AUC of 0.71 (95% CI: 0.53 to 0.90; p = 0.035) for predicting pCR, demonstrating statistically significant discriminative performance. At a fixed specificity of 90 percent (false-positive rate of 10 percent), the model achieved a sensitivity of 46 percent, meaning it identified nearly half of all true pCR patients while maintaining high specificity.
The AI output produced a continuous probability score that spread across the range from 0 to 1, with pCR patients tending to have higher scores than non-pCR patients. The distribution of scores showed meaningful separation between groups, providing a probabilistic output that could be used flexibly depending on clinical decision thresholds rather than a fixed binary classification.
Notably, this AUC of 0.71 is comparable to results from pre-NACT MRI-based AI studies in the published literature. For context, a deep learning model on pre-NACT MRI achieved an AUC of only 0.55 in one study (vs. 0.97 using post-NACT MRI), while multivariate machine learning on pre-NACT MRI achieved AUC of 0.71 in another study, and an I-SPY TRIAL CNN on MRI achieved 0.72. The mammography AI model's AUC of 0.71 is therefore in the same performance range as the best pre-NACT MRI approaches despite using a far more accessible and less expensive imaging modality.
As an informal benchmark, seven experienced radiologists jointly reached an AUC of 0.71 for assessing pCR status from post-NACT mammograms (unpublished data from the NeoDense trial). The AI model achieves a comparable level of discrimination from pre-NACT images, which is a more clinically challenging task than post-treatment assessment where treatment effects are already visible.
The clinical rationale for why mammograms might predict chemotherapy response lies in what these images capture: the tumor's radiological appearance and its relationship to the surrounding breast parenchyma. Tumor features visible on mammograms, including shape, margin characteristics, and density, reflect underlying biology such as tumor grade, proliferation rate, and microenvironmental interactions that influence chemosensitivity.
Beyond tumor appearance, breast parenchymal patterns visible on mammograms, including background density distribution and architectural features, are associated with the breast tissue microenvironment. This microenvironment, including immune infiltration and stromal composition, influences tumor behavior and response to chemotherapy. By analyzing both tumor and reference breast tissue simultaneously, the dual-patch architecture of this AI model may capture some of this contextual biological information.
Multiple predictive factors for NACT response have been identified beyond imaging, including molecular subtype, tumor-infiltrating lymphocytes, Ki-67, and genetic expression profiles. Mammographic AI provides a complementary, non-invasive, and globally accessible source of predictive information that could be combined with these other biomarkers in future multimodal prediction models to improve overall accuracy.
The model's ability to capture useful predictive signal from mammograms acquired over a 14-year period and from three different scanner vendors speaks to a degree of robustness. The pre-processing step of histogram equalization was identified as critical for enabling the detection model to perform consistently across vendor differences, demonstrating the practical importance of image normalization in multi-site AI applications.
The most immediate clinical application of a pre-treatment pCR predictor is enabling response-guided treatment, where the therapeutic strategy is tailored based on predicted likelihood of response before the first cycle begins. Early identification of predicted non-responders could prompt consideration of alternative regimens, clinical trial enrollment, or earlier surgical planning.
In the post-NACT setting, accurate pCR prediction could eventually support less invasive surgical approaches. If imaging combined with minimally invasive procedures could confirm pCR with high confidence, some patients who achieve complete responses might be candidates for surgical de-escalation, a concept under active investigation in ongoing clinical trials. Reliable pre-surgical imaging-based confirmation of pCR is a prerequisite for such approaches.
The primary practical advantage of mammography-based prediction over MRI or PET/CT is the worldwide accessibility of mammography equipment. In many healthcare systems, MRI capacity is limited and costly, while mammography is standard infrastructure. An AI tool that can extract predictive information from routine mammograms could therefore be deployed broadly, including in lower-resource settings where MRI is unavailable.
The authors note that future development should include explainable AI methods such as heat maps that visually highlight which regions of the mammogram the model found most informative. This would make the AI's reasoning transparent to radiologists and oncologists, building trust and enabling clinical integration as a decision-support tool rather than a black box.
This proof-of-concept study established that pre-treatment digital mammograms contain information that AI can leverage to predict pCR to neoadjuvant chemotherapy with an AUC of 0.71, a performance level comparable to the best available pre-treatment MRI-based AI models. The result validates the hypothesis that mammographic patterns related to tumor biology and microenvironment reflect underlying factors that determine chemotherapy sensitivity.
The study's key limitations include the relatively small test set of 53 patients, which limits confidence in the precision of the AUC estimate (wide confidence interval of 0.53 to 0.90), and the inability to perform subgroup analysis by breast cancer molecular subtype due to the AI model's data requirements. Subtype-specific validation is important given the dramatically different pCR rates across subtypes.
The heterogeneous cohort spanning 14 years and multiple scanner vendors is both a strength and a limitation. It demonstrates real-world generalizability but makes it harder to isolate the specific imaging features the model relies on. The long recording period also means some variation in treatment protocols, although the core anthracycline-taxane backbone was consistent throughout.
The next planned phase of this research involves training AI on serial mammograms acquired at multiple time points during NACT, which may improve performance by capturing dynamic changes in tumor appearance as treatment progresses. Combined with explainable AI heat map analysis, this represents a path from proof-of-concept to a clinically interpretable, prospectively validated tool.