Prostate cancer (PCa) is the fifth leading cause of cancer death in men worldwide, with a 13.5% incidence and 6.7% mortality rate. Among the most important prognostic factors is whether cancer has spread to nearby lymph nodes -- a condition called lymph node involvement (LNI). Recurrence of LNI in initially diagnosed patients is reported at 1.3% to 12% and is strongly associated with poor outcomes.
Currently, lymph node status is assessed using imaging tools including MRI, CT, and PSMA PET/CT. However, for MRI and CT, the only criterion used to flag a suspicious lymph node is its size -- specifically a short-axis diameter greater than 1 cm. This size-based approach misses many metastatic nodes that are not yet enlarged, and conversely may flag benign enlarged nodes as suspicious.
To decide which patients need extended pelvic lymph node dissection (ePLND) during radical prostatectomy, clinicians rely on established statistical tools called nomograms -- for example the Briganti and MSKCC nomograms. The European Association of Urology recommends ePLND when estimated lymph node risk exceeds 5%. However, lymph node dissection carries its own surgical risks, including lymphocele, bleeding, nerve damage, and infection.
The emergence of AI-based radiomics -- the extraction of quantitative imaging features from medical scans -- offers a potential path to more accurate, noninvasive lymph node staging. This systematic review compiled and analyzed all studies published through August 2023 examining AI and machine learning approaches to LNI detection and prediction in prostate cancer patients.
Two independent radiologists searched the MEDLINE, PubMed, and Web of Science databases using search terms combining 'prostate,' 'radiomics,' 'machine learning,' 'deep learning,' 'artificial intelligence,' and 'lymph.' No language or date restrictions were applied, and the last search update was completed in August 2023.
From an initial pool of 192 papers, 16 research articles were selected as eligible after applying strict inclusion and exclusion criteria. Included studies had to report on lymph node involvement in prostate cancer patients and use at least one imaging modality (CT, MRI, or PET-CT). Reviews, case reports, editorials, abstracts, and pediatric studies were excluded.
The included studies were classified by imaging modality: eight MRI-based, two CT-based, and six PET-CT-based (using PSMA PET, [18F]DCFPyL PET, and [18F]FMCH PET tracers). Only three of the 16 studies were prospective; the remaining 13 were retrospective in design. The Radiomics Quality Score (RQS) was used to assess the methodological quality of each study.
The review followed PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines to ensure transparent and reproducible reporting. For each included study, data were collected on patient numbers, imaging technique, segmentation software, AI algorithms used, and key performance results.
Several MRI-based AI studies demonstrated strong performance in predicting pelvic lymph node metastasis. The study by Zheng et al. using multiparametric MRI radiomics with a Support Vector Machine (SVM) classifier achieved an impressive AUC of 0.915 in the test set, significantly surpassing existing clinical nomograms (AUCs 0.698 to 0.724). The Bourbonne et al. study using a neural network combining clinical and radiomic features achieved a C-Index of 0.89, also outperforming all major nomograms including Briganti and MSKCC.
The Hou et al. PLNM-Risk calculator -- which integrates radiomics machine learning and deep transfer learning with clinical factors -- achieved AUCs of 0.93, 0.92, and 0.76 in training/validation, internal test, and external test cohorts respectively. Crucially, it spared 59.6% of patients from unnecessary ePLND, compared to only 44.9% for MSKCC and 38.9% for Briganti nomograms, while generating fewer false positives (59.3% vs. 70.1% and 72.7% for the nomograms).
Liu et al. developed a 3D U-Net deep learning algorithm specifically for automated detection and segmentation of pelvic lymph nodes on diffusion-weighted MRI images. This model achieved an AUC of 0.963 in detecting suspicious lymph nodes and a precision, recall, and F1-score of 0.97, 0.98, and 0.97, respectively -- demonstrating that deep learning can reliably automate a task that is typically performed manually by radiologists.
A second Hou et al. study found that a machine learning model integrating clinical features, biopsy data, and MRI findings outperformed the MSKCC nomogram with an AUC of 0.816 (p < 0.001). Below a 15% risk cutoff, the ML model spared more than 50% of patients from unnecessary ePLND while missing fewer than 3% of actual LNI cases -- a clinically meaningful trade-off between surgical risk and diagnostic accuracy.
CT-based radiomics also showed strong potential. Peeken et al. developed a radiomics model using contrast-enhanced CT to predict lymph node metastasis status in patients undergoing PSMA radioguided surgery for recurrent prostate cancer. The combined radiomic model achieved a test-set AUC of 0.95, significantly outperforming conventional CT measurements such as lymph node short diameter (AUC 0.84) and volume (AUC 0.80), demonstrating the clear added value of texture and shape features beyond simple size criteria.
PET-CT-based AI models achieved some of the highest reported performance values in the review. Cysouw et al. applied Random Forest machine learning to [18F]DCFPyL PET radiomics from 76 patients and successfully predicted LNI with an AUC of 0.86, outperforming standard PET metrics. Notably, the PSMA expression captured on PET was linked both to primary tumor histopathology and metastatic tendency, suggesting that radiomics from PSMA PET encodes biologically meaningful information about tumor aggressiveness.
Trägardh et al. developed freely available AI tools using [18F]PSMA-1007 PET-CT that achieved sensitivity on par with nuclear medicine physicians: 79% for detecting prostate tumor/recurrence versus 78% for physicians, 79% for lymph node metastases versus 78%, and 62% for bone metastases versus 59%. In a focused study on pelvic lymph nodes in high-risk patients, a separate AI model achieved 82% sensitivity versus 77% for human readers. The Borrelli et al. AI tool using 18F-choline PET/CT detected more LN lesions than one of the two human readers, while also demonstrating significant associations between the number of detected LN lesions and prostate cancer-specific survival.
A striking observation across all 16 studies is the extreme variability in methodology: different imaging modalities, different radiomic feature sets, different segmentation approaches, and different machine learning algorithms were used, making direct cross-study comparisons difficult. Most studies used traditional radiomics models; only five used deep learning approaches. The most commonly reappearing algorithms across papers were Random Forest and Support Vector Machine (SVM).
Despite this variability, a consistent pattern emerged: combining radiomic features with clinical variables consistently outperformed either alone. The most informative radiomic features were texture-based, first-order statistical, and shape-based features. Notably, Local Binary Pattern (LBP) features -- an innovative addition to standard radiomics -- were tested by Peeken et al. and contributed to their high AUC of 0.95.
One of the most significant methodological limitations identified was the reliance on manual segmentation for feature extraction in most studies. Manual segmentation introduces inter-reader variability, slows workflow, and reduces reproducibility -- all barriers to clinical implementation. Automated segmentation approaches, such as the 3D U-Net models evaluated in several studies, represent an important step toward standardization and scalability.
Other limitations noted include the predominantly retrospective study designs (13 of 16 studies), small sample sizes (five studies enrolled fewer than 100 patients), and the lack of external validation in several works. The review authors conclude that while results consistently exceeded 85% accuracy, standardized methodologies and prospective validation are still needed before these AI tools can be adopted in clinical practice.
The primary clinical motivation for better LNI prediction is the ability to selectively spare patients from extended pelvic lymph node dissection (ePLND) -- a surgical procedure that can cause serious complications including symptomatic lymphocele (in approximately 8% of patients), bleeding, infection, and nerve or vascular damage. If an AI model can reliably identify the roughly 5-15% of patients who actually have lymph node metastases, the majority can safely avoid this unnecessary surgical risk.
Current nomograms like Briganti and MSKCC, while widely used, still result in a large number of unnecessary lymph node dissections because their accuracy is limited. The PLNM-Risk AI calculator reviewed in this study could spare 59.6% of patients from ePLND -- substantially more than the 44.9% spared by MSKCC -- while also reducing false positives, meaning fewer cancer cases would be missed. This represents a clinically meaningful improvement in the quality of surgical decision-making.
AI models that can also automatically segment and monitor treatment response in lymph nodes have additional utility for patients with advanced prostate cancer receiving systemic therapy. The Liu et al. deep learning algorithm achieved 92% accuracy in assessing target lymph node lesion response -- comparable to radiologist performance but without the time and variability associated with manual assessment. This capability is important for following patients over time according to structured response criteria like MET-RADS-P.
This systematic review found that AI-based radiomics models show consistent promise in detecting and predicting lymph node involvement in prostate cancer across MRI, CT, and PET-CT imaging modalities. The best-performing models achieved accuracies consistently above 85% and in several cases substantially outperformed established clinical nomograms such as Briganti and MSKCC.
The key challenge going forward is standardization. The current landscape is fragmented, with no agreed-upon best imaging modality, segmentation approach, or machine learning algorithm for this specific task. The review authors call for larger prospective studies, external validation across institutions, and standardized protocols to enable reproducible comparisons and move toward clinical deployment.
The ultimate goal -- enabling truly personalized surgical planning by accurately predicting which specific patients have lymph node spread before surgery -- remains within reach but requires the field to move beyond its current exploratory phase toward rigorous, standardized, prospectively validated clinical tools. The strong early results reported across these 16 studies provide a compelling rationale for that next step.