Prostate cancer (PC) is the second most frequently diagnosed cancer in men worldwide and the fifth leading cause of cancer-related death. In the United States alone, over 288,000 new cases and nearly 35,000 deaths were projected for 2023, underscoring an urgent need for improved detection tools.
Current standard detection methods include digital rectal examination (DRE), blood tests for prostate-specific antigen (PSA), transrectal ultrasound (TRUS)-guided biopsy, and multi-parametric MRI (mpMRI). Each carries significant limitations: PSA has high false-positive rates, biopsy can miss tumors or over-diagnose harmless ones, and even mpMRI suffers from variability between interpreting radiologists.
The PI-RADS (Prostate Imaging Reporting and Data System) provides a standardized framework for scoring prostate lesions on MRI, but even its most recent version (v2.1) does not meaningfully improve specificity over its predecessor, leading to unnecessary biopsies especially for intermediate-risk PI-RADS 3 lesions. Inter- and intra-observer variability remains a persistent problem.
These gaps create an opening for artificial intelligence (AI) and deep learning (DL) to add value. By automating pattern recognition in MRI images, DL tools can standardize interpretation, reduce reader dependence, and potentially detect tumors that trained radiologists might miss -- all without invasive procedures.
The authors conducted a systematic search of the PubMed/Medline database, searching for studies combining the terms prostate cancer, magnetic resonance imaging, and deep learning, restricted to full-text English articles published between January 2019 and March 2023. All studies had to use 3.0 Tesla MRI scanners for data acquisition.
From 168 records initially identified, 139 were excluded (duplicates, review papers, non-English, non-relevant imaging modalities), leaving 29 original studies involving a combined total of 17,954 participants for inclusion.
Study quality was assessed using two validated tools: the CLAIM (Checklist for Artificial Intelligence in Medical Imaging) checklist, which covers 42 items for AI study quality, and the QUADAS-2 tool for evaluating risk of bias and applicability in diagnostic accuracy studies. The median CLAIM compliance across included studies was 61.9%, indicating moderate but incomplete adherence to reporting standards.
The 29 studies were characterized by year of publication, patient numbers, study design (prospective vs. retrospective), and whether they involved multiple institutions. Most studies (93%) were retrospective, and 62% were conducted at a single center -- limitations that affect how broadly their results can be generalized.
The most commonly used MRI inputs across the 29 studies were T2-weighted imaging (T2WI) and apparent diffusion coefficient (ADC) maps derived from diffusion-weighted imaging (DWI). These sequences together capture both the structural anatomy of the prostate and the functional property of water molecule diffusion, which differs markedly in cancer tissue versus normal tissue.
Deep learning architectures used across studies included a wide range of models: U-Net (a convolutional encoder-decoder widely used for segmentation), ResNet, VGG, AlexNet, Inception-Net, and more specialized custom networks such as Focal-Net, SPCNet, and TrumpetNet. Most were trained to detect, localize, or classify clinically significant prostate cancer (csPC), typically defined as Gleason grade group 2 or higher.
One study used a Variational Network (VN) -- not for detection but for accelerating MRI acquisition, reducing scan time from 11.8 minutes to 3.2 minutes while preserving diagnostic image quality. Another used deep learning reconstruction (DLR) to significantly improve signal-to-noise and contrast-to-noise ratios in DWI images, improving lesion visibility without affecting quantitative ADC values.
Data augmentation techniques -- particularly random rotation for DWI images -- were used in several studies to expand training datasets and reduce overfitting in situations where labeled prostate MRI data was limited. The PROSTATEx public challenge dataset was used in 10 of the 29 studies, reflecting its central role as a benchmark resource in this field.
Multiple DL models demonstrated strong performance in detecting clinically significant prostate cancer. One study achieved an AUC of 0.995 using a 3D convolutional encoder-decoder model with T2WI input. Others consistently reached AUCs in the 0.85 to 0.96 range for csPC detection -- comparable to or exceeding experienced radiologist performance in head-to-head comparisons.
Ten studies directly compared DL outcomes to those of human radiologists. Results showed that DL models often matched or exceeded subspecialist radiologist performance. For example, in one study the AI model and subspecialist radiologist both achieved an AUC of 0.86 for csPC detection, while junior readers achieved only 0.80 and general radiologists 0.83. This suggests DL could help level performance across experience levels.
Several studies found DL outperformed PI-RADS v2.1 scoring in specificity for csPC, meaning fewer false positives and potentially fewer unnecessary biopsies. The PIDL-CS model, combining DL classification with PI-RADS, achieved a significantly higher AUC than PI-RADS alone (AUC 0.881 vs. 0.850, P less than 0.05), validating the benefit of integrating AI into the existing clinical scoring framework.
An integrated risk stratification model called PRISK combined high-throughput MRI features with clinical indicators using stacked ensemble learning. Its performance approached that of invasive biopsy -- achieving 90.4% accuracy on external validation versus 94.2% for biopsy -- positioning it as a promising non-invasive alternative for grading prostate cancer severity.
A particularly noteworthy finding across the review was the limited coverage of transition zone (TZ) prostate cancer -- tumors arising in the central gland region. The TZ accounts for about 30% of all PC cases and is notoriously difficult to evaluate because TZ tumors overlap visually with benign prostatic hyperplasia (BPH) on MRI. Only one study specifically targeted TZ-PC using DL, achieving sensitivity and precision of 0.829 and 0.617 respectively using only ADC as input.
The authors highlight that inter-reader variability in MRI interpretation remains a core unsolved problem, and that DL-based computer-aided diagnosis (DL-CAD) tools offer a meaningful path forward. When DL tools were used as reading aids, they improved diagnostic accuracy and reduced reporting times -- critical benefits for clinical deployment at scale.
Federated learning (FL), an approach where AI models are trained across multiple institutions without sharing raw patient data, was identified as a key technology for multi-center studies. One study showed FL improved generalization performance in lesion classification by 9.5 to 14.8%, with near-complete improvement in lesion segmentation, demonstrating that privacy-preserving multi-site collaboration is both feasible and effective.
Several technical factors were identified as drivers of false positives and false negatives in DL-CAD models: contrast-to-noise ratio (CNR) of lesions, ADC values, lesion diameter, and rectal susceptibility artifacts. Understanding and controlling these variables will be important for improving model reliability in real-world clinical settings where image quality varies across scanners and institutions.
A central ambition expressed by the review authors is to develop DL-assisted mpMRI into a system capable of reducing or replacing invasive prostate biopsy. Biopsy carries real clinical risks including bleeding, infection, and urinary complications, and currently misses or over-detects tumors with meaningful frequency. A non-invasive AI-powered pipeline could change this calculus.
The review calls for prospective, multi-center trials with multi-ethnic cohorts and multi-vendor MRI hardware to validate DL models in representative real-world populations. Current studies are predominantly retrospective and single-center, which limits generalizability and may introduce selection biases not representative of clinical practice.
The review recommends that future DL models be built with radiological and histopathological annotations from diverse patient cohorts, paired with optimized CNN architectures and rigorous external validation. This combination would allow models to learn from the full clinical picture and be tested in genuinely independent settings before deployment.
Ultimately, the authors envision a future where PI-RADS evolves into an AI-assisted system that continuously learns from real-world clinical data, moving away from static scoring rules toward data-driven lesion classification. This would standardize diagnosis globally and reduce the dependency on rare subspecialist radiologist expertise in settings with limited resources.