Residual cancer burden (RCB) is a scoring system that measures how much cancer remains in the breast and lymph nodes after completing a full course of neoadjuvant chemotherapy, assessed through pathological examination of the surgical specimen. RCB 0 means no remaining invasive cancer, equivalent to a pathological complete response. RCB I represents minimal residual disease, RCB II represents moderate residual disease, and RCB III represents extensive residual disease that indicates the cancer was largely resistant to chemotherapy.
The critical clinical problem is that approximately 10 to 35% of breast cancer patients are resistant to NAC, yet this resistance can only be confirmed with certainty after weeks or months of treatment followed by surgical analysis. Patients with drug-resistant tumors, who will ultimately be classified as RCB III, receive toxic chemotherapy that does not benefit them and may delay more effective treatments. An AI system that can predict RCB scores during treatment could enable early identification of resistant patients and prompt changes to treatment strategy.
This study developed the Longitudinal Radiomics and Deep Learning Pipeline (LRadDP), a multitask AI system that analyzes MRI scans taken before NAC and partway through treatment to predict final RCB scores. Rather than relying on a single timepoint, the system captures how tumor characteristics change during treatment, providing a dynamic picture of tumor response that outperforms static pre-treatment analysis.
Validated in 1,048 patients across four institutions, the system achieved AUCs above 0.97 in the primary cohort and 0.90 to 0.94 across three external validation cohorts for both identifying drug-resistant RCB III cases and identifying favorable-response RCB 0 to I cases, demonstrating strong multicenter generalizability.
The current standard for assessing tumor response during NAC is RECIST criteria, which measure changes in the largest diameter of the tumor on imaging scans. While RECIST is simple and widely used, it has well-documented limitations: tumors can change shape without changing size, tumor hypoxia and internal fibrosis can alter imaging appearance without reflecting true cancer cell kill, and pseudoprogression (apparent growth that is actually immune infiltration) can mislead radiologists. The study found that of patients RECIST classified as disease progressors, only 55% were confirmed as RCB III by pathology.
Radiomics is an approach that extracts hundreds or thousands of quantitative features from medical images, including texture, shape, and intensity statistics that are invisible to the human eye. These features can capture subtle patterns in tumor heterogeneity and composition that reflect underlying tumor biology. Radiomics applied to pre-treatment MRI had previously achieved AUCs of 0.76 to 0.82 for predicting treatment response, but these models ignored the dynamic information contained in how the tumor changes during treatment.
Longitudinal radiomics, which computes feature changes between imaging timepoints (called delta-features), provides additional predictive power by measuring tumor evolution directly. A tumor that is destined to be drug-resistant may show subtly different patterns of early change compared to a responding tumor, even before macroscopic size changes are apparent. Prior studies had shown that longitudinal radiomics outperforms single-timepoint analysis for various cancer treatment response prediction tasks.
The study combined radiomics features with deep learning features extracted from MRI images using ResNet-50, creating a richer representation of tumor imaging characteristics than either approach alone. Multitask learning addressed the three-class classification challenge by designing two complementary binary models: one identifying drug-resistant RCB III cases (highest clinical urgency), and one identifying favorable-response RCB 0 to I cases (enabling confident continuation of therapy).
Each patient underwent multiparametric MRI at two timepoints: before NAC began (pre-NAC) and at the mid-point after 3 or 4 completed cycles (mid-NAC). The scans included T2-weighted images, dynamic contrast-enhanced (DCE) sequences (which show how blood vessels deliver contrast agent into the tumor), and diffusion-weighted images (which reflect tissue cellularity and water mobility). Tumor and peritumoral regions were segmented by radiologists using 3D volumetric methods, with a 3D U-Net deep learning model assisting initial segmentation and achieving a Dice coefficient of 0.83 in testing, indicating strong agreement with manual delineations.
From each region of interest, 1,223 radiomics features were extracted per scan, covering intensity histograms, shape descriptors, and multiple gray-level texture matrices. Including the peritumoral zone (5mm around the tumor) alongside the tumor region was a deliberate design choice based on evidence that the tumor microenvironment contributes predictive information. Combined with deep learning features from ResNet-50, each patient had 21,804 radiomics and 6,144 deep learning features across pre-NAC, mid-NAC, and delta-NAC feature sets.
Feature selection was rigorous, using four sequential filters to reduce the initial feature space to a manageable subset of 11 to 15 features per task per timepoint: Mann-Whitney U test to retain features that differed significantly between RCB groups, Spearman correlation analysis to remove redundant features that carried overlapping information, LASSO regression to select features with non-zero predictive coefficients, and the Boruta algorithm with 500 bootstrap repetitions to confirm the most important features. All retained features were verified to have intraclass correlation coefficients above 0.75, ensuring measurement reproducibility.
The final stacking pipeline combined submodels trained on each individual feature set (pre-NAC, mid-NAC, and delta-NAC) with significant clinical variables into a meta-model using SVM. Model I classified RCB III versus RCB 0 to II, incorporating HER2 status as a known strong predictor of drug resistance. Model II classified RCB 0 to I versus RCB II to III, incorporating clinical T stage, estrogen receptor status, and HER2 status, reflecting the biological reality that both imaging and molecular receptor characteristics jointly determine treatment sensitivity.
Model I, which identifies drug-resistant RCB III cases, achieved an AUC of 0.975 and accuracy of 96.12% in the primary cohort. Across the three external validation cohorts, it maintained AUCs of 0.922, 0.936, and 0.910, with sensitivities of 96.60%, 94.25%, and 94.86%. This high sensitivity for RCB III detection is clinically critical: missing drug-resistant patients (false negatives) means continuing futile chemotherapy, so sensitivity must be maximized even if some false positives are accepted.
Model II, which identifies favorable-response RCB 0 to I cases, achieved AUCs of 0.976 and accuracy of 92.54% in the primary cohort, with external validation AUCs of 0.903, 0.910, and 0.918. High sensitivities of 88.5%, 90.24%, and 88.28% across external cohorts enable clinicians to confidently identify responders who should receive the full recommended course of treatment and are candidates for breast-conserving surgery after NAC.
Compared to the RECIST imaging standard, LRadDP showed substantially higher sensitivity for identifying both RCB III and RCB 0 to I cases. For RCB III detection, LRadDP achieved sensitivities of 70.59% to 80.78% across external cohorts versus 20.41% to 25.00% for RECIST, representing a three to four fold improvement in identifying drug-resistant patients. This dramatic difference illustrates how much prognostic information the human eye misses when interpreting standard imaging criteria.
Subgroup analysis confirmed consistent performance across cancer molecular subtypes (HER2-positive, hormone receptor positive/HER2-negative, and triple-negative) and clinical T stages (cT1-2, cT3, cT4). Performance did not significantly vary by subtype or disease extent, indicating the model captures universal imaging signatures of treatment response rather than relying primarily on molecular subtype as a proxy for response, which would limit its added value over existing clinical information.
The two-model design of LRadDP maps directly onto the two most pressing clinical decisions during NAC. Model I addresses the question of which patients are unlikely to benefit from continuing chemotherapy: a high RCB III prediction score, available after just 3 to 4 cycles, could prompt a clinician to consider switching to alternative chemotherapy regimens, enrolling the patient in a clinical trial of salvage therapies, or planning earlier surgery before further tumor progression occurs.
Model II addresses the opposite question: which patients are responding well enough to continue and potentially qualify for less extensive surgery. Breast-conserving surgery rather than mastectomy is associated with better quality of life and is achievable only when the tumor has shrunk substantially. Early confirmation that a patient is on track for RCB 0 to I response can support shared decision-making between the oncologist and patient about surgical planning and chemotherapy continuation.
The inclusion of peritumoral regions in the model, rather than analyzing only the visible tumor mass, reflects a growing understanding that the microenvironment surrounding the tumor plays an important biological role in treatment response. Immune cell infiltration, blood vessel density, and tissue stiffness in the peritumor zone all reflect the body's response to cancer and to the drug, and these characteristics are measurable through MRI texture features even when not visible by standard radiological interpretation.
The noninvasive nature of MRI-based prediction is a major practical advantage over tissue-based approaches. Repeated biopsies during chemotherapy are technically challenging, associated with patient discomfort and procedural risk, and subject to sampling bias from tumor heterogeneity. MRI can image the entire tumor and its environment repeatedly throughout treatment at low patient burden, making longitudinal monitoring at scale feasible in routine clinical practice.
The study's retrospective design is its primary methodological limitation. Although the four-institution enrollment and three external validation cohorts provide stronger evidence than a single-center study, the patient selection, imaging protocols, and chemotherapy regimens were not standardized prospectively. Patients with incomplete MRI data, non-standard regimens, or distant metastases were excluded, which may have selected for a more typical patient population than would be encountered in a fully prospective clinical deployment.
The molecular subtype imbalance in the dataset (HER2-positive patients were the largest group at 45.71%) may have influenced model performance, since HER2 status is one of the strongest known predictors of NAC response. The models incorporated HER2 status as a clinical feature, which is appropriate since this information is routinely available, but it also means that part of the model's predictive power may derive from molecular subtype rather than pure imaging features. Validation in cohorts with different subtype distributions would further establish generalizability.
The semi-automated tumor segmentation process, which required radiologist review and correction of the 3D U-Net segmentation outputs, adds a time burden that would need to be eliminated or dramatically reduced for routine clinical deployment. A fully automated end-to-end system from raw MRI acquisition through RCB prediction without radiologist intervention remains the goal for practical clinical use, and the authors identify this as a key development priority.
Despite these limitations, the study provides the most comprehensive multicenter validation to date of longitudinal MRI-based RCB prediction, with external cohort AUCs consistently above 0.90. Prospective clinical trials where treatment decisions are guided by LRadDP predictions are the appropriate next step to determine whether the model's accuracy translates into genuine improvements in patient outcomes, including survival, surgical outcomes, and reduction of futile chemotherapy exposure.
This study demonstrates that combining longitudinal MRI radiomics and deep learning features from pre- and mid-treatment scans provides substantially more accurate prediction of final residual cancer burden than single-timepoint imaging or standard RECIST criteria alone. The LRadDP system achieves this through an analytically rigorous pipeline: automated tumor segmentation, comprehensive 3D feature extraction including peritumoral zones, multi-step feature selection, and stacking of single-modality models with clinical variables.
The three to four fold improvement in sensitivity for RCB III detection over RECIST is potentially the most clinically impactful finding. Correctly identifying drug-resistant patients during rather than after chemotherapy provides a window for individualized treatment adaptation, which is the practical goal of precision oncology in breast cancer NAC management.
Validation across four institutions in China with consistent performance across different patient populations, scanner types, and chemotherapy regimens supports the generalizability of the model's learned imaging signatures. These are not institution-specific technical artifacts but appear to represent biologically meaningful features of tumor response and resistance that are measurable across different clinical settings.
As breast cancer NAC protocols continue to evolve and as MRI becomes more routinely available in cancer centers worldwide, noninvasive AI tools like LRadDP represent a practical path toward making treatment decisions more responsive to individual tumor biology, potentially sparing patients from ineffective therapy and improving overall outcomes across the breast cancer population receiving neoadjuvant chemotherapy.