Neoadjuvant chemotherapy (NAC) is commonly prescribed to reduce tumor burden prior to breast cancer surgery, improve surgical outcomes, and allow breast conservation in otherwise unresectable tumors. Pathological complete response (pCR), defined as the absence of any residual invasive disease, is a key indicator of treatment success. Patients who achieve pCR are more likely to be candidates for breast-conserving surgery and have longer progression-free and overall survival.
The ability to predict which patients will respond to NAC early in the treatment course is clinically important because it could help minimize unnecessary toxic chemotherapy and allow modification of regimens mid-treatment. A major challenge is the lack of reliable methods to assess efficacy early in the NAC course. Radiological prediction of pCR using MRI is a desirable non-invasive alternative to pathology, as it can assess the entire breast without being limited to a biopsied area.
Machine learning is increasingly used in radiology and medicine to identify relationships among complex data elements. Convolutional Neural Networks (CNNs) are particularly well suited for medical image analysis because they can operate on whole images without requiring radiologists to manually contour tumors. This feature-agnostic approach contrasts with supervised ML methods that rely on extracted features like tumor volume and radiomic characteristics.
This systematic review focuses on deep learning methods that use whole-breast MRI images without annotation or tumor segmentation to predict pCR. The review compares image types used (DCE or post-contrast), whether pre-training, transfer learning, or data augmentation was employed, and whether models incorporate molecular subtypes, multiple treatment time points, and multi-institutional data.
The literature search used PubMed and Google Scholar with keywords combining breast MRI, deep learning, and pathological complete response. From 79 initially identified studies, 66 were removed through screening, leaving 13 studies for the final systematic review.
Thirteen studies were included in the systematic review, employing various CNN architectures including VGG13, VGG16, AlexNet, ResNet-50, ResNeXt50, and custom CNN designs. The studies used different MRI image types including dynamic contrast-enhanced MRI (DCE), contrast-enhanced (CE), and T2-weighted images. Sample sizes ranged from 25 to 536 patients across the studies.
Performance metrics varied substantially across studies. Area under the curve (AUC) values ranged from 0.72 to 0.97, with accuracies spanning from 72.5% to 92.3%. The best-performing model by Qu et al. achieved an AUC of 0.97 using a 2D CNN with 12 channels combining 6 DCE phases from both pre- and post-NAC images along with molecular subtypes including ER, PR, and HER2 status.
Braman et al. demonstrated the feasibility of deep learning-based prediction using data from multiple sites, achieving an AUC of 0.93 and accuracy of 86.7% with a multiphasic CNN applied to pre-NAC DCE MR images of HER2-positive breast cancer patients.
Several studies demonstrated that incorporating non-imaging data such as molecular subtypes, demographics, and clinical variables alongside MRI data improves prediction performance. HER2-positive and triple-negative tumors achieved significantly higher rates of pCR compared to luminal A subtype tumors. Hormone receptor positive and HER2-positive cancers have specific receptors that can be targeted with directed therapy.
Duanmu et al. showed that integrating both MRI and non-MRI data outperformed models using imaging data alone or traditional concatenation approaches. Their integrated CNN model evaluated 3D DCE whole images at multiple treatment timepoints incorporating molecular subtype and demographic data, achieving an accuracy of 0.81 and AUC of 0.83 for PCR prediction. The study highlighted that epigenetic factors and changes in molecular status during treatment progression should also be considered.
Most studies used single post-contrast MRI images, though some investigated multiple DCE dynamics. Research found that among multiple DCE dynamics, the first post-contrast image combined with pre-contrast data was most effective for predicting response to therapy. The pre-contrast image serves as a baseline, while early post-contrast images capture important vascular and physiological tumor characteristics.
While DCE data is most widely used, other MRI data types including T2-weighted imaging, diffusion imaging, and fat-water imaging could also inform pCR prediction. Background parenchymal enhancement by MRI is informative of cancer risk, recurrence, and outcomes but has not yet been adequately explored. CNN predictive models using multiple DCE images were shown to be superior to those using individual dynamics alone.
Three broad challenges must be overcome before deep learning can achieve mainstream clinical applications: generalizability, interpretability, and ethical/legal concerns. Training datasets need to be not only large but also diverse to avoid or minimize bias. Publicly available high-quality clinical data for testing pCR predictive models are currently limited. Federated learning offers a potential solution by training models on multiple local datasets without data sharing across institutions.
Deep learning findings are difficult to interpret due to the complex nature of the calculations and the many features from which conclusions are drawn. Tools such as Shapley values and heatmaps can improve interpretability. There are also ethical and legal uncertainties regarding responsibility in the case of incorrect diagnoses, and the black-box nature of deep learning systems makes it challenging to determine how models arrive at predictions.
Computer-assisted diagnosis systems with deep learning AI can assist radiologists in detecting breast cancers and increase diagnostic efficiency. Deep learning can potentially serve as a secondary or concurrent reader, increasing accuracy while decreasing interpretation time. Automated triaging to prioritize scans with urgent findings is another practical application that could reduce radiologist burnout.
The ability to predict treatment responses early during NAC allows for individualized treatment and precision medicine. This could be particularly beneficial in underserved areas with lower socioeconomic populations who currently suffer from worse breast cancer outcomes. Automated preprocessing, segmentation, detection, and classification of lesions may reduce unnecessary biopsies and surgeries by predicting the behavior of precancerous lesions.
Deep learning prediction of pCR could have a central role in breast cancer management and make a positive impact on patient care. However, studies in the current literature generally do not have large enough data to achieve broad generalizability, and the potential of deep learning in predicting pCR is not yet fully realized. Rigorous comparison of different models using the same datasets is also needed.
Future deep learning studies should intelligently integrate multiple types of MRI data including DCE, T2-weighted, fat-water imaging, and diffusion MRI, along with imaging data at multiple treatment time points and molecular subtypes, demographics, and genetic data. In addition to predicting pCR, deep learning models can be used to predict residual cancer burden, progression-free survival, risk of recurrence, and overall survival. Broad adoption requires further clinical validation, improved reliability, generalizability, and interpretability.