MRI-based AI models for post-neoadjuvant surgery personalization in breast cancer

Lancet Reg Health West Pac 2025 MRI Analysis 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Personalizing Surgery Using AI and MRI After Chemotherapy

Neoadjuvant systemic therapy -- chemotherapy, targeted therapy, or immunotherapy given before surgery -- has become standard care for locally advanced breast cancer. One of the most important goals is determining, before surgery, whether a patient has achieved a pathological complete response (pCR), meaning all invasive cancer cells have been eliminated. This information is critical for deciding how extensive surgery should be.

Traditional surgical planning relies on physical examination, imaging, and biopsy, but these methods are imperfect and often lead to overly aggressive surgery even in patients who responded fully to chemotherapy. MRI-based artificial intelligence models offer a promising way to non-invasively assess treatment response and guide decisions such as whether to proceed with breast-conserving surgery, axillary lymph node dissection, or observation.

This narrative review, published in Lancet Regional Health -- Western Pacific in 2025, systematically examined 51 published studies on AI models that use MRI imaging to personalize surgical planning after neoadjuvant therapy. The review focused on the Western Pacific region, where the majority of this research originates, and evaluated both the clinical performance and methodological quality of the evidence base.

TL;DR: This review examines 51 studies on MRI-based AI models that predict treatment response after neoadjuvant therapy to guide breast cancer surgical decisions.
Pages 2-3
Study Scope: 51 Papers, Five Surgical Endpoints

The review identified 51 eligible studies published primarily from China (88%), South Korea, and Japan. The vast majority were retrospective (94%) and conducted at single institutions (80%), with a median sample size of only 152 patients (interquartile range 114 to 329). This reflects the early stage of the field, where research has clustered in academic centers with access to MRI archives and digital pathology records.

Studies were categorized by their primary endpoint: pCR prediction was the most common focus (70.59% of studies), followed by axillary lymph node response, tumor regression pattern classification, residual cancer burden, and disease-free survival. Each endpoint corresponds to a distinct surgical decision -- from mastectomy versus breast-conserving surgery, to whether lymph node dissection can be safely omitted.

MRI protocols varied across studies: 54.9% used dynamic contrast-enhanced MRI (DCE-MRI) alone, while 45.1% incorporated multiparametric MRI combining DCE-MRI with functional sequences such as diffusion-weighted imaging (DWI) and T2-weighted imaging. Study quality was assessed using the Radiomics Quality Score (RQS), a 16-item tool that evaluates methodological rigor in AI imaging studies.

TL;DR: 51 predominantly retrospective, single-center studies from China used DCE-MRI or multiparametric MRI to predict five surgical endpoints, most commonly pathological complete response.
Pages 3-5
pCR Prediction Performance Across Studies

Across 28 studies reporting pCR prediction in validation cohorts, the pooled sensitivity was 0.78 (95% CI 0.72 to 0.83) and specificity was 0.87 (95% CI 0.83 to 0.91). AUC values ranged from 0.71 to 0.97 depending on MRI protocol, AI methodology, and patient population. These figures indicate that current models are generally better at ruling in pCR than ruling it out -- meaning they occasionally miss patients who still have residual disease.

A key finding was that multitemporal radiomics -- also called delta-radiomics -- consistently outperformed models using a single MRI timepoint. By extracting features from MRI scans taken before, during, and after neoadjuvant therapy and analyzing how they change over time, these models achieved pooled sensitivity of 0.83 and specificity of 0.88. Capturing treatment-induced changes in tumor texture, volume, and vascularity provided substantially more predictive information than any single scan alone.

Models integrating multiparametric MRI (combining structural and functional sequences) also outperformed DCE-MRI alone. Adding diffusion restriction measurements and perfusion parameters reflecting tumor cellularity and vascularity enriched the feature space and improved discrimination between complete and incomplete responders. Across all approaches, deep learning models incorporating attention mechanisms and convolutional architectures achieved higher AUCs than traditional radiomics with handcrafted features.

TL;DR: AI models predicted pCR with pooled sensitivity 0.78 and specificity 0.87, with delta-radiomics and multiparametric MRI approaches achieving the strongest performance.
Pages 5-6
Beyond pCR: Lymph Node and Tumor Regression Prediction

Several studies addressed axillary lymph node response -- whether cancerous lymph nodes have cleared after chemotherapy -- a surgical decision point that determines whether full axillary dissection is necessary. AI models predicted lymph node pathological complete response with sensitivity ranging from 0.71 to 0.88 and specificity from 0.80 to 0.94, with accuracy between 0.83 and 0.88 across studies. This suggests AI could help identify patients who might safely avoid full lymph node dissection.

Tumor regression pattern classification -- distinguishing concentric shrinkage from heterogeneous fragmentation -- is relevant because different regression patterns require different surgical margins. AI models trained to classify regression patterns achieved AUC values of 0.83 to 0.94, providing potentially useful pre-surgical information about how the tumor has responded spatially. Fragmented regression is particularly challenging surgically because residual foci may be difficult to delineate by imaging.

A smaller number of studies predicted residual cancer burden (RCB) -- a continuous measure of residual disease -- and disease-free survival after surgery. These endpoints are clinically important because RCB class directly determines whether patients receive additional adjuvant therapy such as capecitabine or olaparib. AI models for RCB and survival prediction showed promising initial results but were based on very few studies with limited external validation.

TL;DR: AI models predicted axillary lymph node clearance with accuracy up to 0.88 and tumor regression patterns with AUC up to 0.94, supporting multiple distinct surgical decisions.
Pages 6-7
The Breast-Conserving Surgery Gap in the Western Pacific

A central clinical motivation for this work is the persistently low rate of breast-conserving surgery (BCS) in China and other Western Pacific countries, even among patients who achieve pCR after chemotherapy. Studies show that Chinese patients with confirmed pCR still choose mastectomy at rates far higher than their Western counterparts, reflecting cultural preferences, limited surgeon confidence in imaging-guided decisions, and patient anxiety about leaving any breast tissue.

AI models that reliably predict pCR before surgery could shift this dynamic by giving surgeons and patients quantitative, imaging-based confidence that conservative surgery is appropriate. For patients predicted to have complete tumor elimination, the model output serves as a decision-support tool that supplements standard imaging and clinical judgment. For patients predicted to have residual disease, it guides more aggressive surgical planning or neoadjuvant escalation strategies.

The clinical stakes extend beyond cosmetic outcomes. Patients who undergo unnecessary mastectomy face greater surgical morbidity, longer recovery, and significant quality-of-life consequences. AI-guided surgical personalization thus addresses both oncological safety and patient wellbeing -- ensuring that treatment intensity matches the actual extent of residual disease rather than a worst-case assumption.

TL;DR: AI pCR prediction models could address the under-utilization of breast-conserving surgery in China by providing quantitative imaging-based confidence to guide surgical decisions.
Pages 7-8
Radiomics Methodology and Quality Assessment

Radiomics is a computational approach that extracts hundreds or thousands of quantitative features -- such as texture, shape, intensity distribution, and spatial heterogeneity -- from medical images. In the studies reviewed, radiomics features were extracted from manually or semi-automatically segmented tumor regions on MRI scans. These features capture imaging phenotypes that are invisible to the naked eye but may reflect underlying tumor biology.

Delta-radiomics extends this approach by computing feature differences between multiple MRI timepoints -- typically before, during (mid-treatment), and after neoadjuvant therapy. The change in radiomics features over time captures dynamic tumor response to treatment, including regression of vascularity, reduction in cellularity, and changes in tissue texture that precede measurable size reduction. These longitudinal features carry more predictive information than the post-treatment scan alone.

Study quality was assessed using the Radiomics Quality Score (RQS), which evaluates 16 methodological criteria including feature reproducibility testing, external validation, clinical utility assessment, and prospective study design. The median RQS across all 51 studies was only 12 out of 36 -- indicating widespread methodological limitations. Most studies lacked feature stability analysis, prospective registration, calibration statistics, and formal clinical decision analyses.

TL;DR: Delta-radiomics capturing longitudinal MRI feature changes outperformed static approaches, but the median study quality score of 12/36 reveals widespread methodological gaps.
Pages 8-9
Racial Bias and Generalizability Concerns

A critical finding in the review was evidence of racial bias in AI model performance. One study initially developed and validated on a German cohort (88.6% white patients) showed substantially degraded performance when applied to a Chinese patient cohort. This performance drop reflects the fact that breast cancer biology, imaging characteristics, and even MRI acquisition protocols differ across populations -- and models trained on one demographic may not generalize to others.

This finding has direct implications for the Western Pacific region, where the majority of patients are of East Asian descent but where reference models and publicly available training datasets are often derived from predominantly white, Western populations. If AI tools developed primarily in Europe or North America are adopted without revalidation on Asian patient cohorts, they risk systematic underperformance for the very populations they are intended to serve.

Addressing generalizability requires deliberate study design: prospective multi-center trials enrolling patients across diverse racial and ethnic backgrounds, standardized MRI acquisition protocols to reduce scanner variability, and federated learning approaches that allow model training across institutions without requiring data sharing. The review authors emphasize that equity in AI model performance must be a design requirement, not an afterthought.

TL;DR: AI models trained on Western populations showed degraded performance on Asian cohorts, highlighting racial bias as a critical challenge for Western Pacific clinical adoption.
Pages 9-10
What Must Change Before Clinical Deployment

The review identifies several structural barriers to clinical translation. Nearly all studies were retrospective, meaning they analyzed historical data with no prospective patient enrollment or pre-registered analysis plans. Retrospective designs are vulnerable to selection bias, data leakage, and overfitting -- problems that inflate apparent model performance and create false confidence in results that may not hold in real clinical settings.

Methodological gaps identified across studies include absent calibration statistics (which measure whether predicted probabilities match actual outcomes), no formal clinical utility analysis (such as decision curve analysis to show whether acting on model predictions improves outcomes), and limited feature robustness testing (which assesses whether features are stable across scanners, segmentations, and acquisition settings). Without these, it is impossible to know whether a model is actually ready for clinical use.

The authors call for a new generation of studies: prospective, multi-center trials with pre-registered analysis plans, standardized MRI protocols, racial and geographic diversity, external validation on independent cohorts, and formal integration into surgical decision-making workflows. Only through this standard of evidence can MRI-based AI models move from research tools to clinical instruments that reliably guide individualized surgical care for breast cancer patients across the Western Pacific region and beyond.

TL;DR: Clinical deployment requires prospective multi-center validation, standardized protocols, calibration and utility analyses, and deliberate design for racial diversity -- none of which characterizes the current evidence base.
Citation: Open Access, 2025. Available at: PMC12121432.