Radical prostatectomy (RP) -- surgical removal of the prostate gland -- is the primary treatment for localized or locally advanced prostate cancer, and it achieves excellent long-term survival when the disease is caught in time. However, up to 50% of patients experience cancer recurrence after surgery, detected first as rising PSA levels in the blood.
This rising PSA after surgery is called biochemical recurrence (BCR). It is the earliest warning sign of cancer returning and is strongly linked to later spread to distant organs and to cancer-related death. Catching patients at high risk for BCR early allows doctors to start additional treatments before visible metastasis appears.
Known risk factors for BCR include high preoperative PSA levels, aggressive tumor grades (high Gleason score), cancer spreading beyond the prostate capsule (extracapsular extension), invasion of the seminal vesicles, and cancer cells at the surgical margin after removal. However, even among patients with these high-risk features, outcomes vary widely -- indicating that current clinical tools miss important biological information.
This study asks whether deep learning analysis of preoperative multiparametric MRI (mpMRI) -- the detailed prostate imaging performed before surgery -- can more accurately predict which patients will experience long-term BCR-free survival after radical prostatectomy.
The study enrolled 437 patients with prostate cancer who underwent radical prostatectomy at Samsung Medical Center in Korea between 2008 and 2009, with preoperative mpMRI performed before each surgery. Patients were followed up every three months for the first two years, every six months through year five, and annually thereafter. During a median follow-up of 61 months, 110 patients (25.2%) experienced BCR.
This Korean patient cohort is important for a specific reason: Asian men, particularly Koreans, tend to present with more adverse pathological features at diagnosis -- higher Gleason scores, more advanced tumor stages -- than Western populations. Standard Western clinical prediction tools have been shown to perform suboptimally in Asian patients, motivating development of a population-specific model.
The unique aspect of this dataset is that routine preoperative prostate mpMRI before radical prostatectomy was standard practice at this institution since 2008, which is uncommon even today and has allowed accumulation of over 15 years of mpMRI data linked to long-term clinical outcomes. This makes it one of the few centers capable of building and testing mpMRI-based deep learning models for BCR prediction.
BCR was defined using the widely accepted threshold of PSA rising to 0.2 ng/mL or above on two consecutive measurements after surgery -- a strict, validated criterion for detecting recurrence.
The researchers built and compared five different prediction models using the same patients. The clinical model (CM) used six standard pathological variables: patient age, preoperative PSA, extracapsular extension, seminal vesicle invasion, positive surgical margins, and Gleason grade. This represents what doctors currently use to estimate recurrence risk.
The radiomics model (RM-Multi) extracted 72 mathematically defined features (shape, intensity, and texture) from manually contoured tumor regions on three MRI sequences (T2-weighted, ADC maps, and contrast-enhanced images). This is the conventional imaging analysis approach that requires an expert to outline tumors slice by slice.
The deep learning model (DLM) used a pretrained neural network called EfficientNet-B0 -- originally trained on millions of natural images -- and fine-tuned it on the prostate MRI data. Rather than requiring manual tumor outlining, the model automatically learned features from image patches centered on the tumor, capturing not just the tumor but also the surrounding tumor microenvironment (TME). 320 features per MRI sequence (960 total from three sequences) were extracted from intermediate network layers.
The two remaining models combined imaging with clinical data: CRM-Multi merged the radiomics features with the six clinical variables, and CDLM merged the deep learning features with the six clinical variables. All models used Cox proportional hazards regression with LASSO regularization and were evaluated using stratified fivefold cross-validation to prevent overfitting.
The combined clinical-deep learning model (CDLM) achieved the best performance of all five models, with a hazard ratio (HR) of 7.72, C-index of 0.89, and integrated time-dependent AUC (iAUC) of 0.93 on the test set. An HR of 7.72 means patients classified as high-risk were 7.72 times more likely to experience BCR than low-risk patients -- a clinically very meaningful separation.
The standalone deep learning model came second (HR 4.37, C-index 0.74, iAUC 0.77), outperforming the radiomics model (HR 2.67, C-index 0.66, iAUC 0.68) by a substantial margin. The clinical model alone performed strongly (HR 5.82, C-index 0.81, iAUC 0.85), reflecting that the six pathological variables still carry important prognostic information.
Notably, the clinical-radiomics combination (CRM-Multi) failed to outperform the clinical model alone (HR 5.44 versus 5.82), suggesting that handcrafted radiomics features did not add useful information beyond standard clinical data. This contrasts with the deep learning features, which provided genuinely complementary information when added to clinical variables.
Grad-CAM visualization -- a technique that shows which image regions the deep learning model focuses on -- confirmed that the model was correctly attending to areas near the primary tumor when making its predictions, providing reassurance that the network was making biologically meaningful decisions rather than exploiting unrelated image artifacts.
The superiority of deep learning over radiomics stems from fundamental architectural differences. The EfficientNet-B0 network contains 17 layers of convolution that progressively extract features from fine-grained to abstract, capturing nonlinear relationships between image characteristics that fixed handcrafted radiomics features cannot represent.
Deep learning features are also more flexible -- they adjust to the specific characteristics of the local data during training, rather than being predefined. This means deep features can capture disease-relevant patterns specific to this Korean patient population, while radiomics features calculated by fixed formulas cannot adapt in this way.
The image patch approach in this study also included information about the tumor microenvironment (TME) -- the tissue immediately surrounding the visible tumor boundary. The TME contains immune cells, blood vessels, and structural changes that influence tumor behavior and are invisible to standard radiomics approaches but captured in the deep learning feature space.
Using the pretrained EfficientNet-B0 (pre-trained on ImageNet's vast image database) allowed the model to work well with a modest dataset of 437 patients -- a key advantage since prostate MRI datasets with long follow-up remain relatively rare and expensive to collect.
The most direct clinical application is using this model's pre-surgery MRI risk score to guide postoperative monitoring intensity and adjuvant therapy decisions. Patients classified as high-risk by the CDLM model could be monitored more frequently after surgery and considered for early adjuvant radiation or hormonal therapy, which evidence suggests works best when started before visible recurrence.
Current standard tools like the CAPRA-S scoring system have known limitations -- they require lymph node dissection data that many patients do not have, and they rely on biopsy Gleason scores that can underestimate true tumor aggressiveness. This deep learning approach, using surgical pathology results combined with preoperative MRI, avoids some of these pitfalls.
The fact that the model uses routine preoperative mpMRI -- imaging already performed for surgical planning -- means no additional tests or costs are required to generate the risk prediction. This makes clinical translation relatively straightforward once validation is complete.
Future integration with PSMA-PET imaging -- a newer, highly sensitive scan for prostate cancer -- could provide additional complementary information about tumor biology and micrometastatic spread, potentially further improving prediction accuracy beyond what MRI alone can achieve.
This study demonstrates that combining deep learning features from preoperative multiparametric MRI with standard clinical pathology data creates the most powerful currently available tool for predicting long-term biochemical recurrence-free survival after radical prostatectomy, outperforming both conventional radiomics and clinical data alone.
The hazard ratio of 7.72 and C-index of 0.89 for the combined model represent state-of-the-art performance for this prediction task and exceed the results of comparable studies using radiomics or simpler deep learning approaches.
The key limitations are that this is a single-center retrospective study using imaging data from 2008-2009, with tumor delineation performed by one expert. While cross-validation was rigorous, external validation in prospective multicenter cohorts is essential before clinical deployment.
The authors envision future work combining this approach with multi-omics data (genomics, proteomics) and incorporating newer imaging modalities like PSMA-PET to build even more comprehensive and robust risk models for the individualized management of prostate cancer after surgery.