Evaluation of a Cascaded Deep Learning-based Algorithm for Prostate Lesion Detection at Biparametric MRI

Radiology 2024 Deep Learning 7 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
The Challenge of Consistent Prostate MRI Interpretation

Multiparametric MRI (mpMRI) has become the standard imaging approach for detecting prostate cancer before biopsy. It allows radiologists to identify suspicious lesions and guide targeted biopsies that sample only the most concerning areas of the prostate rather than random cores.

Despite standardized reporting guidelines (PI-RADS version 2.1), prostate MRI interpretation remains subject to significant variation between readers. Different radiologists often assign different categories to the same lesion, leading to inconsistent decisions about whether and where to biopsy.

AI algorithms based on deep learning can potentially standardize this interpretation process by providing an objective, automated assessment of every scan. However, most published AI models have been evaluated only on small, retrospective datasets from the same institution where they were developed -- not on large prospective samples reflecting real clinical practice.

This study evaluated a previously developed cascaded deep learning AI model for lesion detection and segmentation in a large prospective clinical sample of 658 patients, comparing its performance directly against an expert genitourinary radiologist with over 15 years of experience and over 1,000 scans per year.

TL;DR: Prostate MRI interpretation varies significantly between radiologists; this study evaluated a cascaded deep learning AI model against an expert radiologist in a large prospective patient sample.
Pages 2-3
How the Cascaded Deep Learning Algorithm Works

The AI model is a cascaded deep learning algorithm that processes biparametric MRI input -- T2-weighted images, high-b-value diffusion-weighted images, and apparent diffusion coefficient (ADC) maps. It does not use dynamic contrast-enhanced imaging, making it compatible with the growing trend toward contrast-free prostate MRI protocols.

The algorithm operates in two stages: first it automatically segments the prostate gland itself, then within that prostate region it detects and delineates intraprostatic lesions suspicious for cancer. This cascade structure reduces the search space and focuses computational resources on the relevant anatomy.

The model incorporates a benign prostatic hyperplasia (BPH) filter to reduce false positives in the transition zone, where BPH nodules can mimic cancer on imaging. This helps achieve a balance between detecting all cancer foci and minimizing spurious detections that would reduce clinical utility.

The algorithm outputs both binary prediction masks (definitive lesion outlines) and probability maps (showing the confidence level across the prostate). These outputs could be exported directly into clinical PACS workstations to overlay on the radiologist's viewing screen during routine clinical reading.

TL;DR: The cascaded AI algorithm uses biparametric MRI inputs to first segment the whole prostate, then identify and delineate intraprostatic lesions, outputting both detection masks and probability maps.
Pages 2-4
Study Design and Patient Population

The study included 658 male patients (median age 67 years) from a prospective clinical registry at the National Institutes of Health, enrolled between April 2019 and September 2022. All had undergone mpMRI followed by biopsy -- either US-guided systematic biopsy alone (for PI-RADS 1 scans) or combined systematic and MRI/US fusion-guided targeted biopsy (for PI-RADS 2-5 lesions).

The radiologist reference standard was provided by a single expert genitourinary radiologist who prospectively scored all lesions using PI-RADS version 2.1 and volumetrically contoured the whole prostate and each lesion. A second radiologist from a different institution independently evaluated a random 51-patient subset to assess inter-reader variability as a benchmark for the AI comparison.

The study's gold standard for cancer presence was histopathologic analysis of biopsy samples, with clinically significant prostate cancer (csPCa) defined as International Society of Urological Pathology (ISUP) grade group 2 or higher -- the threshold that typically changes clinical management from active surveillance to active treatment.

AI detection performance was assessed at two levels: lesion level (did the algorithm detect each individual radiologist-defined lesion?) and participant level (did the algorithm correctly identify which patients had cancer overall?). The Dice similarity coefficient (DSC) was used to measure how closely the AI's lesion outlines matched those drawn by the radiologist.

TL;DR: Six hundred fifty-eight patients with prospective PI-RADS evaluation and biopsy results were used to evaluate AI performance at both the individual lesion level and the patient level.
Pages 5-6
Patient-Level Detection Performance vs. Radiologist

At the patient level, the AI algorithm detected 96% (282 of 294) of all participants with clinically significant prostate cancer, while the expert radiologist detected 98% (287 of 294). This difference was not statistically significant (p = 0.23), indicating the AI performed comparably to the expert radiologist for the most important clinical task: identifying which patients truly need treatment.

For all prostate cancer regardless of grade, the AI detected 92% of affected patients versus 93% for the radiologist (p = 0.75). Performance increased with cancer grade: the AI detected 96% of ISUP grade group 2, 96% of grade group 3, 95% of grade group 4, and 98% of grade group 5 participants -- showing the algorithm was most reliable for the most aggressive cancers.

The algorithm also detected lesions in 76 of 99 participants assigned PI-RADS category 1 by the radiologist (considered negative scans). Among these, systematic biopsy confirmed cancer in 25% and csPCa in 8%, including 6 of 7 cases where a negative scan actually harbored clinically significant cancer -- suggesting the AI may detect some cancers that expert radiologists miss.

The positive predictive value (PPV) of the AI was significantly lower than the radiologist at PI-RADS 2 or above threshold: 47% versus 51% for csPCa (p < 0.001). This means the AI generated more false-positive detections per scan, a trade-off for its comparable sensitivity.

TL;DR: The AI detected 96% of patients with clinically significant prostate cancer versus 98% for the expert radiologist, a non-significant difference, but had lower positive predictive value due to more false-positive detections.
Pages 6-7
Lesion-Level Detection and Segmentation Performance

At the individual lesion level, the AI algorithm demonstrated a sensitivity of 55% (569 of 1029 radiologist-defined lesions detected) and a positive predictive value of 57%. The mean number of false-positive lesion predictions per patient was 0.61, meaning the algorithm flagged on average less than one additional lesion per scan beyond those the radiologist identified.

Detection performance was strongly correlated with PI-RADS category: the algorithm detected 30% of PI-RADS category 2 lesions, 42% of category 3, 61% of category 4, and 92% of PI-RADS category 5 lesions. This performance gradient reflects both the algorithm's strength with high-confidence lesions and its limitation in detecting subtle, lower-category findings.

For lesion segmentation accuracy, the overall Dice similarity coefficient (DSC) was 0.29 -- meaning the AI's drawn lesion boundaries partially but imperfectly matched those drawn by the radiologist. DSC improved dramatically with lesion aggressiveness: 0.14 for benign findings versus 0.58 for PI-RADS category 5 lesions and 0.50 for ISUP grade group 5 cancers.

Analysis revealed that the AI tended to segment the central core of lesions rather than their full extent, while radiologists incorporated peripheral features and adjacent tissue characteristics. This systematic difference explains the moderate overall DSC even when the AI correctly identified lesion location.

TL;DR: Lesion-level sensitivity was 55% overall but reached 92% for PI-RADS 5 lesions; segmentation accuracy improved with lesion aggressiveness, reaching DSC 0.58 for the highest-risk lesions.
Pages 7-9
Clinical Applications and Deployment Potential

The strongest clinical application for this AI model is as a patient triage tool -- identifying which patients have cancer requiring biopsy at the participant level, where it matched radiologist performance. For institutions with limited access to expert genitourinary radiologists, an AI-based pre-screen could help ensure high-risk patients are not missed.

The algorithm's excellent performance on PI-RADS 5 lesions (92% detection, DSC 0.58) makes it particularly useful for biopsy planning in the highest-risk group. Automated lesion segmentation at this category could reduce the manual annotation workload for MRI-guided biopsy preparation and radiation therapy planning.

As a second reader deployed within clinical PACS workstations, the AI could alert radiologists to lesions or patients they might have categorized as negative, using its probability maps and detection masks as an additional reference. This is especially valuable for radiologists with less prostate MRI specialization than the expert used in this study.

The detection of clinically significant cancer in patients with PI-RADS category 1 scans (considered negative by the radiologist) by the AI -- including 6 of 7 true csPCa-positive cases -- raises the interesting possibility that AI could identify a small but important subset of cancer patients who would otherwise be reassured and not biopsied.

TL;DR: The AI's strongest clinical role is as a patient triage tool and second reader, particularly valuable for PI-RADS 5 biopsy planning and for identifying the rare cases of cancer missed by negative radiologist assessments.
Pages 8-9
Limitations and Path to Broader Validation

All MRI scans were evaluated and annotated by a single expert radiologist, meaning the AI was compared against one person's interpretation rather than a consensus ground truth. While this is actually a stringent test (matching a leading expert), it does not measure how the AI compares to typical community radiologists who have less prostate MRI experience.

The AI model does not incorporate dynamic contrast-enhanced imaging because acquisition was inconsistent in the training dataset. While biparametric MRI (without contrast) is increasingly accepted as adequate for prostate evaluation, this limits evaluation of whether adding contrast sequences would improve detection, particularly for transition zone tumors where DWI is less reliable.

The study was conducted at a single academic institution with high patient cancer prevalence (45% with csPCa) and consistent MRI acquisition parameters. Performance may differ in community settings with more heterogeneous scanning protocols, different equipment, and lower cancer prevalence -- conditions that typically reduce AI performance compared to development-site results.

The authors conclude that a large prospective multicenter study comparing this algorithm against radiologists at different experience levels is needed to fully establish generalizability and clinical utility. Additionally, an interventional reader study where radiologists interact with the AI tool -- rather than comparing them independently -- would more accurately predict real-world clinical impact.

TL;DR: Single-center evaluation against one expert radiologist limits generalizability; multicenter validation across diverse imaging settings and experience levels is needed before widespread clinical adoption.
Citation: Open Access, . Available at: PMC11140533.