Extraprostatic extension (EPE) occurs when prostate cancer grows through the outer capsule of the prostate gland into the surrounding tissue. It is present in approximately 20-30% of prostate cancer cases and is an independent predictor of poor outcomes: patients with EPE face higher rates of positive surgical margins, biochemical recurrence, and distant metastasis following radical prostatectomy.
Accurately knowing whether EPE is present before surgery profoundly changes the surgical plan. In EPE-negative patients, the surgeon can preserve the adjacent neurovascular bundle -- the network of nerves controlling erection -- significantly reducing the risk of postoperative erectile dysfunction and urinary incontinence while maintaining cancer control. In EPE-positive patients, a wider resection is required to minimize the risk of leaving cancer cells behind.
Current preoperative EPE assessment relies on multiparametric MRI (mpMRI), which assesses tumor capsular contact length, capsular irregularity, and signs of invasion. However, radiologist accuracy for EPE detection using standard MRI criteria varies widely, with AUC values typically between 0.60 and 0.75. Many EPE cases involve only microscopic extension that lacks clear visible features on conventional imaging, making them particularly difficult to detect.
This study developed and validated a deep learning model combining mpMRI and 18F-PSMA-PET/CT imaging to predict EPE, and evaluated whether AI assistance could improve the accuracy of radiologist assessments beyond what either tool achieves alone.
The retrospective study enrolled 197 patients from Center 1 (training and internal testing) and 36 patients from Center 2 (external validation), all with pathologically confirmed prostate cancer who underwent radical prostatectomy. All patients had both mpMRI and 18F-PSMA-1007 PET/CT scans performed within 4 weeks before surgery, and no patient had received neoadjuvant therapy that could alter imaging appearance.
Each patient's prostate region of interest was carefully delineated on all imaging modalities, then resampled to uniform dimensions (64x64x8 voxels) to create consistent inputs for the deep learning network. Pixel values were normalized to a 0-1 range, and data augmentation using random affine transformations was applied to improve model generalization and reduce overfitting.
The deep learning model used a ResNet-50 architecture -- a 50-layer residual convolutional neural network. Residual networks address a common problem in deep networks (vanishing gradients) by adding skip connections that allow learning to bypass certain layers, enabling very deep networks to train effectively. For multimodal fusion, all five imaging channels (ADC, DWI, T2WI from MRI; CT and PET from PSMA-PET/CT) were concatenated into a single multi-channel input at the network's entry point.
A separate DL-assisted radiologist model was also constructed. A radiologist with over 10 years of experience independently scored each case using the established EPE-grade system (0-3 points based on MRI findings). The AI model's prediction was then used to adjust the radiologist's score by plus or minus one point, depending on whether the model predicted EPE-positive or negative. This hybrid approach was designed to test whether AI could systematically improve clinical assessments.
In the internal test group (Center 1), the mpMRI-only model achieved an AUC of 0.76 for EPE prediction, while the PSMA-PET/CT-only model achieved 0.77. The combined multimodal model integrating both achieved a significantly superior AUC of 0.82, with 79% sensitivity and 78% specificity. These results demonstrate that the two imaging modalities provide complementary rather than redundant information about EPE risk.
In the external validation cohort (Center 2), consistent results were observed: mpMRI alone achieved AUC 0.75, PSMA-PET/CT alone achieved 0.77, and the multimodal model achieved AUC 0.81 with 94% sensitivity. The high sensitivity in the external set is particularly important clinically, as missing a true EPE case (false negative) leads to inadequate surgical margins and higher recurrence risk.
Comparing the imaging modalities, mpMRI showed higher specificity (78% vs. 71% for PET/CT in Center 1) while PSMA-PET/CT showed higher sensitivity (76% vs. 65% for mpMRI). This complementary performance profile explains why combining them improved overall classification: mpMRI catches the clear anatomic EPE features while PSMA-PET/CT detects metabolic activity that may signal microscopic extension not yet visible on structural imaging.
The DL-assisted radiologist model showed the most dramatic improvement over baseline: radiologist EPE-grade scoring alone achieved an AUC of only 0.64 at Center 1, while the DL-assisted scoring model raised this to AUC 0.82 -- an improvement of 0.18 AUC points (p less than 0.001). The external validation at Center 2 showed an even larger improvement, from AUC 0.68 to 0.87.
Beyond statistical performance metrics, the study evaluated clinical net benefit using decision curve analysis (DCA). This method assesses whether a model actually improves clinical decisions by comparing the net benefit of using the model against two default strategies: treating all patients as EPE-positive (very wide resection for everyone) or treating none as EPE-positive (nerve-sparing for everyone).
When the threshold EPE probability exceeded 0.6, the DL-assisted EPE-grade score model provided a significantly higher net benefit than either the radiologist EPE-grade score alone or the standalone DL model. This means that in clinical practice, using the AI-assisted approach would lead to better overall outcomes -- more correct decisions about when to preserve vs. resect the neurovascular bundle -- than either tool used independently.
The improvement in sensitivity (catching more true EPE cases) came at a small cost in specificity (slightly more false positives -- patients predicted EPE-positive who were actually EPE-negative). For surgical planning, this tradeoff is generally acceptable: an unnecessary wider resection carries some quality-of-life consequences, but missing actual EPE carries a higher risk of cancer recurrence and metastasis.
This study also confirms that preoperative EPE assessment is genuinely challenging: even experienced radiologists using structured EPE-grade scoring achieved only AUC 0.64-0.68 in this cohort. The lower-than-expected baseline performance may reflect a higher proportion of microscopic EPE cases in this Asian population, where subtle features are harder to detect on conventional MRI than the clear invasion patterns studied in training cohorts from other populations.
The complementary value of PSMA-PET/CT reflects fundamental differences in what each modality measures. mpMRI assesses tumor anatomy and tissue microstructure -- it detects EPE by looking for capsular contact, irregularity, and gross extension. PSMA-PET/CT measures metabolic and molecular activity -- it detects cells expressing PSMA protein, which can be present beyond the visible capsule even when structural imaging appears normal.
This means PSMA-PET may detect microscopic EPE that lacks clear anatomic signs on MRI. Conversely, MRI may provide high specificity for clear anatomic EPE features that appear as non-specific uptake on PET. The deep learning model can learn how to weight these complementary signals appropriately across thousands of image voxels, capturing patterns that human radiologists may miss or weight differently.
Comparison with prior work supports the model's competitive performance. A previous DL model using mpMRI alone (PAGNet) achieved AUC 0.86 internally but dropped to 0.73 in external validation. The current multimodal model maintained 0.81-0.82 AUC across both internal and external sites, suggesting better generalizability -- an important indicator of real-world clinical reliability.
This is reported as the first study to combine mpMRI and PSMA-PET/CT in a deep learning model specifically for EPE prediction, and to validate it with external data. While the sample sizes are modest (197 and 36 patients), the consistent performance across institutions provides initial proof of concept for a multimodal AI approach to this challenging clinical problem.
This study demonstrates that a multimodal deep learning model combining mpMRI and PSMA-PET/CT can improve preoperative EPE assessment, with AI assistance raising radiologist AUC from 0.64 to 0.82 internally and 0.68 to 0.87 externally. This translates directly into more informed surgical planning: better identification of patients who need wider resection versus those who can safely have nerve-sparing surgery.
The model's performance as a decision support tool -- adjusting radiologist scores rather than replacing radiologist judgment -- is a pragmatic clinical design. This approach preserves the clinician's role and expertise while systematically correcting for the cases where AI detects patterns not visible on standard interpretation. It reflects a 'human-in-the-loop' AI model appropriate for high-stakes surgical decisions.
Limitations include the relatively small sample sizes, the retrospective design, and the single-institution origin of the external validation cohort. Future research should validate the model in larger, prospective multicenter studies including diverse patient populations. Adding interpretability tools such as activation heatmaps would help radiologists understand which imaging regions drove each prediction, further building clinical trust and facilitating workflow integration.