Prostate cancer is the second most common cancer in men worldwide and a leading cause of cancer-related death. Accurate, early detection is critical for good outcomes. Magnetic resonance imaging (MRI) has become the leading imaging tool for prostate cancer because it offers a complete, operator-independent view of the entire prostate gland, from base to apex and including areas poorly assessed by older methods like digital rectal examination.
Standard prostate MRI, called multiparametric MRI (mpMRI), includes three types of imaging: T2-weighted (T2w) imaging for anatomy, diffusion-weighted imaging (DWI) for cell density, and dynamic contrast-enhanced (DCE) imaging which requires intravenous injection of a gadolinium contrast agent. Reporting is standardized using the PI-RADS (Prostate Imaging Reporting and Data System) scoring framework.
DCE imaging is increasingly being dropped from clinical protocols because it adds risk (gadolinium exposure), cost, and complexity while often adding limited diagnostic value over T2w and DWI alone. The resulting two-sequence approach is called biparametric MRI (bpMRI). Current evidence suggests bpMRI performs comparably to mpMRI for most patients, especially those presenting for the first time.
Because bpMRI uses just two standardized sequences, it is an attractive target for artificial intelligence (AI) post-processing. The standardized image types remove many of the variables that complicate AI analysis of mpMRI, making bpMRI a natural foundation for automated cancer detection and grading tools.
Radiomics is a method that extracts large numbers of quantitative features from medical images, such as texture, shape, and intensity statistics, that may be invisible to the human eye. These features can reflect tumor biology and are often fed into machine learning (ML) models for cancer detection or grading.
Machine learning algorithms are trained on labeled data to recognize patterns and make predictions. In prostate imaging, ML is typically applied after a radiologist manually outlines a suspicious region, meaning human input is still required. This manual step introduces variability and limits automation.
Deep learning (DL), particularly convolutional neural networks (CNNs), can learn directly from images without manual feature extraction. CNNs process images through multiple layers that progressively identify more complex patterns, from edges to anatomical zones to tumor characteristics. State-of-the-art CNN architectures such as U-Net and FocalNet can automatically segment the prostate and identify suspicious lesions.
The key advantage of DL over radiomics-based ML is that DL generates its own features optimized for the specific classification problem, rather than relying on predefined, hand-crafted features. However, DL is often criticized as a black box, making its clinical decision rationale harder to interpret and explain to physicians and patients.
This systematic review queried PubMed in August 2021 using search terms combining prostate MRI with machine learning, deep learning, and radiomics. To capture current techniques only, the search was restricted to publications from 2019 to 2021. From an initial 95 publications, 29 met inclusion criteria and were analyzed in detail.
The 29 included studies collectively analyzed 7,466 patients across a wide range of study sizes, from 25 to 834 patients. Imaging was performed at 3 Tesla (3T) in 93% of studies, with the remainder at 1.5T. The most common imaging input was T2w imaging combined with ADC maps and DWI, used in over half the studies.
Thirteen studies used classic ML techniques (44.8%), 14 used DL techniques (48.2%), and 2 combined both approaches. Clinical data extracted from each study included patient counts, age, AI technique used, MRI sequences, sensitivity, specificity, accuracy, and AUC (area under the curve) as a performance measure.
A notable limitation of direct comparison across studies was that seven of the 29 relied on the same publicly available ProstateX dataset from Radboud University, meaning results are not fully independent. Study endpoints also varied widely, from Gleason score prediction to PI-RADS classification, extracapsular extension, and biochemical recurrence.
Across the 29 studies, both ML and DL approaches demonstrated that AI can detect and grade prostate cancer at a level comparable to trained radiologists. No consistent advantage was found for either ML or DL over the other in overall performance, and results were broadly similar between the two paradigms.
For distinguishing clinically insignificant from clinically significant prostate cancer (csPCA), AUC values ranged widely from about 0.71 to 0.99. The highest AUC values (near 0.999) came from studies that excluded small or poorly delineable lesions, suggesting study design strongly influences reported performance.
Six studies directly compared AI to human radiologists reading the same cases. In all of these, AI performance was statistically similar to radiologist performance, with no study showing a clear, significant advantage for either. For specific tasks such as detecting cancer in the transitional zone or identifying clinically significant cancer in PI-RADS 3 lesions, AI approaches showed potential superiority over radiologist PI-RADS assessment.
One study found that automated CNN-based PI-RADS scoring matched radiologist-assigned scores with no statistically significant difference for PIRADS 3 to 5 lesions. Another demonstrated that AI predicted Gleason score better than human radiologists in both the peripheral and transitional zones, which would be particularly valuable for monitoring patients on active surveillance.
Several studies went beyond simple cancer detection to predict future outcomes from bpMRI images. One large study of 459 patients used radiomics to achieve an AUC of 0.905 for distinguishing benign from malignant tissue, 0.728 for predicting extracapsular extension (ECE), and 0.766 for predicting positive surgical margins (PSM) after surgery.
A DL-based study confirmed an identical AUC of 0.728 for ECE prediction using CNN architectures, reinforcing that preoperative MRI carries information about surgical outcomes that AI can reliably extract. Predicting these outcomes before surgery could help patients and surgeons understand the likelihood of complete cancer removal.
One of the most forward-looking studies predicted biochemical recurrence (BCR), the return of detectable PSA after treatment, from T2w images alone, achieving a C-index of 0.802, which outperformed the Gleason scoring system (C-index 0.583). This multicenter study of 485 patients demonstrated that imaging features can predict disease trajectory beyond the current surgical specimen.
A study in patients under active surveillance found that DL detection accuracy for csPCA improved with tumor volume, reaching sensitivity of 94% and specificity of 74% for tumors above 0.5 cc. This is particularly relevant since active surveillance patients need repeat assessments to detect disease progression early.
Despite promising results, several technical challenges remain. MRI lacks standardized intensity scales, meaning that the same tissue appears with different signal values across different scanners, scanner settings, and even repeat scans on the same machine. Normalizing MRI intensities for AI processing is a necessary but error-prone preprocessing step.
Most studies relied on data from a single institution, making it difficult to assess whether results generalize across different scanners, patient populations, and imaging protocols. Only a handful of the 29 studies were multicentric, limiting confidence in broad applicability. The heavy reliance on the ProstateX dataset across seven studies further limits independence of results.
Even the definition of biparametric MRI is inconsistent across AI studies. Some used T2w with ADC, others T2w with high b-value DWI, others all three DWI-derived images. Without standardized input definitions, cross-study comparisons and regulatory approval of AI tools become difficult. A European initiative is now working toward standardizing AI-based prostate MRI tools.
DL algorithms, particularly CNNs, remain a clinical black box: they can produce accurate results without explaining how. This lack of interpretability creates physician hesitancy and regulatory barriers. Commercialization by larger companies with resources for regulatory certification may be the path to overcoming this barrier and making AI prostate tools widely available.
This systematic review demonstrates that AI-assisted bpMRI is a promising and maturing field. Detection of clinically significant prostate cancer and differentiation from benign tissue using ML and DL is feasible, with performance comparable to that of experienced radiologists across most studies.
Early approaches required radiologists to manually delineate suspicious regions before AI analysis, but newer algorithms, particularly CNN-based deep learning models, can now automatically segment the entire prostate gland and identify lesions without human contouring. This represents an important step toward fully automated workflows.
The most consistent finding is that no single AI approach has yet been established as the gold standard, and there is currently no AI method in routine clinical use for prostate MRI. However, the trend toward CNN-based methods and the appearance of a first commercial AI tool for prostate MRI signal that clinical translation is near.
Priority areas for future work include larger multicenter studies, standardized definitions of bpMRI sequences for AI input, external validation datasets, and greater interpretability of deep learning decisions. These advances are necessary before AI can be routinely trusted as a diagnostic partner in prostate cancer management.