This PRISMA-DTA systematic review assessed whether AI-based Computer-Aided Diagnosis (CAD) systems can match or exceed the performance of radiologists for prostate cancer diagnosis on MRI. Unlike general reviews, it required each included study to compare AI directly to human reader performance on the same patient cases -- a more stringent inclusion criterion than most prior reviews.
Twenty-seven studies published between 2013 and 2021 met the selection criteria. Studies were categorized into three groups based on task: ROI Classification (ROI-C) -- classifying pre-identified suspicious regions -- included 16 studies; Lesion Localization and Classification (LL&C) -- simultaneously finding and characterizing lesions -- included 10 studies; and Patient Classification (PAT-C) -- classifying the whole patient as cancer-positive or negative -- included 1 study.
Study quality was assessed using the QUADAS-2 tool, a validated framework for evaluating diagnostic test accuracy studies. QUADAS-2 evaluates risk of bias in patient selection, index test design, reference standard quality, and study flow. The review also applied additional quality criteria specific to AI studies in radiology.
The median cohort size across all 27 studies was only 98 patients -- far smaller than what is typically needed to power reliable head-to-head comparisons between AI and expert radiologists. Most studies used retrospective single-center data, and no study included a prospective evaluation design.
Across all 27 studies, the most common MRI protocol was multiparametric MRI (mpMRI), combining T2-weighted imaging, diffusion-weighted imaging (DWI), and dynamic contrast-enhanced (DCE) sequences. One study used biparametric MRI (bpMRI) without contrast. Most studies used 3 Tesla field strength scanners from single vendors, which limits generalizability.
The most common CAD algorithms were traditional machine learning (logistic regression, random forests, support vector machines) operating on radiomic features extracted from manually drawn regions of interest. Convolutional neural networks (CNN) were used in five studies, while one study evaluated the commercially available Watson Elementary system.
The endpoint definition for clinically significant cancer varied across studies, most commonly Gleason score (GS) 3+4 or greater (equivalent to Grade Group 2). However, some studies used GS 3+3 (Grade Group 1) as the threshold, which is considered clinically insignificant disease and should be managed with active surveillance. This inconsistency complicates comparisons between studies.
Dataset preparation varied considerably: six studies used random splitting, five used temporal splitting (training on older cases, testing on newer ones), five used leave-one-patient-out cross-validation, four used independent internal test cohorts, four used external data from different institutions, and one study did not describe its separation strategy at all.
For ROI Classification, the high-quality study by Dinh et al. showed that their CAD system at a 95% sensitivity threshold achieved 40% specificity versus only 9% for radiologist Likert scoring at threshold 3 -- meaning CAD would reduce unnecessary biopsies by 31% while missing no clinically significant cancers on the internal test set. However, when the same system was validated externally by Transin et al., CAD sensitivity dropped to 89% versus 97% for radiologists, reversing the advantage.
For Lesion Localization and Classification, results were mixed. The study by Zhu et al. showed that when radiologists could use CAD probability maps as an additional reference (while still scoring freely), per-patient sensitivity increased from 84% to 93% and specificity from 56% to 66%. However, in studies where radiologists were restricted to only accepting or rejecting CAD-highlighted areas, some studies showed sensitivity drops, confirming that the mode of human-AI interaction critically affects outcomes.
Studies using external validation cohorts systematically showed smaller AI advantages or no advantage at all compared to studies using internal validation. The four external validation studies (Gaur, Mehralivand, Transin, Thon) found CAD performance either inferior to or on par with radiologists, while several internal validation studies found CAD superior -- a pattern consistent with AI models overfitting to their training data distribution.
The only PAT-C study (Deniffel et al.) used a CNN to classify patients as having or not having csPCa. The CNN achieved higher per-patient sensitivity (100%) and specificity (52%) than PI-RADS-based radiologist reading, but because the decision threshold was determined from the test set rather than pre-specified, this performance likely overestimates prospective accuracy.
The QUADAS-2 assessment revealed that while most studies had low risk of bias for patient selection and study flow, risk of bias was high or unclear in several domains. Six studies determined their optimal classification threshold using test set data -- a methodological flaw that inflates reported sensitivity and specificity and makes prospective performance unpredictable.
Eight studies had high applicability concerns regarding patient cohort composition: they included patients from biopsy-proven populations who subsequently underwent radical prostatectomy, which introduces spectrum bias because this subgroup has a higher rate of significant cancer than an unselected biopsy population, making results overly optimistic for real-world screening contexts.
Only four of 27 studies used external test sets for final performance reporting. Only three studies statistically justified their sample sizes. Only three studies made their CAD systems publicly available. These gaps in transparency and rigor prevent independent replication and make it impossible to determine whether any individual system is ready for clinical use.
Only two studies incorporated non-imaging clinical data (PSA density, digital rectal examination) alongside MRI features. Multiple studies have shown that combining MRI findings with PSA density substantially improves diagnostic accuracy, yet most AI systems in this review ignored this readily available clinical information.
The review identifies a fundamental mismatch between how most AI studies are designed and what clinicians actually need to know. Stand-alone AUC measurements on retrospective single-center data tell researchers whether an AI can distinguish cancer from non-cancer under ideal conditions -- they do not tell clinicians whether using the AI in a real workflow will improve patient outcomes.
The generalization gap between internal and external validation performance is the most clinically consequential finding. AI systems that appear to outperform radiologists on their own training data routinely fail to maintain this advantage when tested on patients from different institutions. This gap likely reflects differences in scanner hardware, acquisition parameters, patient demographics, and local biopsy practices.
The mode of human-AI interaction matters enormously for clinical outcomes. Studies where radiologists could use AI as a second opinion (Zhu et al.) showed benefit; studies where radiologists were constrained to only review AI-identified areas sometimes showed harm (reduced sensitivity). The optimal clinical workflow -- where to place AI in the diagnostic pathway and how to present its output -- requires specific prospective studies that none of the included papers addressed.
The review strongly recommends prospective clinical trials with pre-specified outcomes, external multicenter validation, and well-defined intended clinical use as the next step for AI in prostate MRI diagnosis. The current evidence base, while extensive, is insufficient to justify widespread clinical deployment of any specific CAD system.
The review's core conclusion is that no current evidence supports deployment of AI for initial prostate cancer diagnosis on MRI. Every study that showed AI superiority over radiologists used internal validation only; every study using external data found AI equal to or worse than radiologists. This pattern is a warning sign, not a minor caveat.
Future studies must be designed around a specific clinical use from the outset. ROI-C systems designed to reduce unnecessary biopsies for indeterminate PI-RADS 3 lesions require different study designs than LL&C systems intended to assist less experienced radiologists in detecting all suspicious areas. Mixing use cases undermines the interpretability of results.
The community urgently needs large, publicly available, well-labelled prostate MRI datasets from multiple institutions and scanner vendors. Without standardized benchmarking datasets analogous to ImageNet for natural images, it is impossible to fairly compare AI algorithms from different research groups or track progress over time.
Prospective evaluation that simulates clinical deployment -- including real-time integration of CAD into radiology workflows, tracking changes in radiologist decision-making, and measuring patient outcomes such as biopsy rates and cancer detection rates -- will provide the definitive evidence needed to determine whether AI truly improves prostate cancer diagnosis in practice.