The Role of Radiomic Analysis and Different Machine Learning Models in Prostate Cancer Diagnosis

J Imaging 2025 Machine Learning 6 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
The Diagnostic Gap: Why Radiomics and Machine Learning Are Needed

Prostate cancer is the second most common cancer in men worldwide, and accurate grading is essential for treatment decisions. Currently, diagnosis begins with PSA (prostate-specific antigen) blood testing and, increasingly, multiparametric MRI (mpMRI) before biopsy. However, PSA has poor specificity -- only about 18% of men with elevated PSA are actually diagnosed with cancer, meaning the vast majority undergo unnecessary biopsies with associated risks of bleeding, infection, and urinary complications.

MRI using the PI-RADS v2 scoring system achieves diagnostic AUC values around 0.89 in expert hands, but performance degrades substantially outside specialized centers due to limited availability of expert prostate radiologists and significant interobserver variability in image interpretation. The large spectrum of MRI acquisition parameters across different scanners and institutions further reduces reproducibility.

Radiomic analysis addresses this problem by extracting quantitative features from MRI images -- including texture, shape, and signal intensity patterns -- that are not easily detected by human observers. When combined with machine learning (ML), these features can be used to build objective classifiers that grade tumors without relying on subjective radiologist interpretation. This approach offers the potential to standardize prostate cancer grading across centers and imaging systems.

This study evaluated seven ML algorithms applied to biparametric MRI (bpMRI) -- a streamlined protocol using only T2-weighted (T2W) and diffusion-weighted (DWI) sequences, without the contrast injection required for full mpMRI. The study specifically tested whether adding the PSA blood test result as a third input alongside imaging radiomic features significantly improved classification performance across multiple clinical grading questions.

TL;DR: PSA testing and MRI alone have significant diagnostic limitations for prostate cancer grading, and this study tested whether combining radiomic imaging features with PSA using seven machine learning models could improve objective classification.
Pages 2-4
Study Design: Multiple Centers, Multiple MRI Systems, Seven Algorithms

The study enrolled 214 men with elevated PSA or clinical symptoms who underwent bpMRI at three different imaging centers equipped with four MRI systems -- two operating at 3 Tesla and two at 1.5 Tesla field strength. All patients subsequently underwent transrectal ultrasound-guided (TRUS) biopsy to confirm lesion type and Gleason grade. The use of multiple centers and scanner types reflects real-world clinical heterogeneity, though it also introduces variability that can challenge model performance.

Lesion segmentation was performed manually on T2-weighted images by an expert radiologist with 10 years of experience using ITK-SNAP software, following PI-RADSv2.1 guidelines. Images underwent standardized preprocessing including bias field correction, intensity normalization, and resampling to a 1x1x1 mm isotropic voxel size. Radiomic features were extracted using Pyradiomics v1.3.0, yielding 101 features per imaging sequence: 14 shape-based features, 18 first-order statistics, and 69 texture features from gray-level co-occurrence, size zone, run length, and dependence matrices.

Feature selection used the Gini index algorithm to identify the most informative features, reducing high dimensionality that could degrade classifier performance. The seven ML classifiers evaluated were: k-Nearest Neighbors (k-NN), Naive Bayes (NB), Logistic Regression (LR), Support Vector Machine (SVM), Decision Tree (DT), Random Forest (RF), and Neural Network (NN). Each model was tuned to its best-performing hyperparameter configuration as established in prior literature.

Three clinical classification questions were evaluated: (1) benign versus malignant lesions (Gleason score 6 or below versus above 6), regardless of lesion location; (2) low-risk versus intermediate-risk lesions in the peripheral zone (PZ) (ISUP 1 versus ISUP 2 and 3); and (3) distinguishing ISUP 2 from ISUP 3 within intermediate-risk PZ lesions. These three tasks reflect clinically meaningful decision points: whether to treat at all, and how aggressively to treat.

TL;DR: Seven machine learning algorithms were trained on radiomic features from biparametric MRI at four different scanner systems across three centers, with and without PSA, to answer three clinical grading questions.
Pages 5-7
PSA Transforms Model Performance: The Key Finding

The most striking finding was the dramatic impact of including PSA as an input variable alongside imaging radiomic features. When using T2W and DWI radiomic features alone -- whether individually or combined -- model performance was limited, with AUC values ranging from 0.703 to 0.807 for the benign versus malignant classification task. This range is broadly consistent with published literature using imaging-only inputs and has uncertain clinical utility.

Adding PSA to the imaging features produced a dramatic improvement across all models and all three classification tasks. For benign versus malignant classification, AUC values jumped to: Neural Network 0.992, SVM 0.957, Decision Tree 0.953, Random Forest 0.946, Logistic Regression 0.884, k-NN 0.868, and Naive Bayes 0.830. For the harder intermediate ISUP grading tasks in the peripheral zone, PSA similarly transformed performance, with several models reaching AUC values of 0.995-1.000 for ISUP 2 versus ISUP 3 discrimination.

Model performance was consistent across the two evaluation strategies employed. Neural Network achieved AUC 0.992 in cross-validation and 0.936 in the held-out validation set, demonstrating robust generalization. SVM showed a perfect AUC of 1.000 on the held-out test set. Random Forest and Decision Tree showed modest drops from cross-validation to held-out testing (0.946 to 0.814 and 0.953 to 0.929, respectively), while simpler models like k-NN, Logistic Regression, and Naive Bayes showed moderate but stable performance across evaluation methods.

The DeLong statistical test confirmed that the performance differences between algorithms were statistically significant for most pairwise comparisons -- meaning the differences in AUC reflect genuine algorithmic capability differences rather than chance variation. Neural Network and SVM consistently outperformed simpler models like Naive Bayes (p less than 0.001 for most comparisons), particularly for the more challenging ISUP grading tasks.

TL;DR: Adding PSA to radiomic MRI features dramatically improved all seven machine learning models, with Neural Network achieving AUC 0.992 and SVM reaching perfect AUC 1.000 in held-out testing for benign-malignant classification.
Pages 8-10
Grading Intermediate-Risk Cancers: A Persistent Challenge

Distinguishing intermediate-risk cancers from each other -- specifically ISUP grade 2 (Gleason 3+4) from ISUP grade 3 (Gleason 4+3) within peripheral zone lesions -- is clinically important because these two categories carry different prognoses and guide different treatment decisions, yet they are difficult to separate on standard imaging. Without PSA, most models achieved AUC values around 0.6-0.85 for this task, with high variability across algorithms.

Adding PSA brought substantial improvement for intermediate grading as well: Random Forest reached AUC 0.995, Logistic Regression 0.972, and SVM 1.000. This pattern held for the task of distinguishing low-risk (ISUP 1) from intermediate-risk (ISUP 2 and 3) lesions, where Logistic Regression and SVM achieved AUC 1.000. These results are higher than most published literature, which typically reports AUCs of 0.73-0.85 for these grading tasks using imaging-only inputs.

The advantage of complex models (Neural Network, SVM) over simpler ones (Logistic Regression, k-NN, Naive Bayes) was most pronounced for grading tasks rather than simple benign-malignant separation. This likely reflects the ability of more powerful algorithms to capture nonlinear interactions between radiomic texture features and PSA levels that simpler linear classifiers cannot detect.

The study also confirms that lesion location matters: the literature shows that AI models generally perform better in the peripheral zone than the transition zone of the prostate, likely because MRI contrast between cancerous and normal tissue is stronger in the peripheral zone. Mixing peripheral and transition zone lesions in the same model without accounting for location is a documented source of performance variability across published studies.

TL;DR: Grading within intermediate-risk prostate cancers remains the hardest task for ML models, but adding PSA and using complex algorithms like Neural Network and SVM substantially improved performance even for these subtle distinctions.
Pages 10-11
Why Results Vary Across Studies: Protocol Differences and DWI b-Values

A key finding from comparing this study's results against published literature is that imaging acquisition parameters -- particularly DWI b-values -- are a major driver of model performance variability. Studies using higher b-values (b = 2000 s/mm2) on single 3T scanners tend to report higher AUC values than this study, which used b-values of at least 1000 s/mm2 across both 1.5T and 3T systems. Higher b-values provide greater contrast between cancerous tissue and surrounding normal tissue, providing more discriminative radiomic features.

The multi-scanner, multi-center design of this study -- though more representative of real-world clinical practice -- inherently introduces variability that reduces classification accuracy compared to idealized single-scanner datasets. Preprocessing steps including bias correction, intensity normalization, and resampling partially mitigate scanner differences, but cannot fully eliminate them. This explains why this study's imaging-only AUCs (0.703-0.807) fall in the lower range of published literature.

The dominant positive role of PSA in this study contrasts with some published studies where PSA added little or even slightly decreased model performance when combined with imaging features. These conflicting results likely reflect differences in dataset composition -- particularly the proportion of transition zone lesions, which have higher background PSA from benign prostatic hyperplasia, making PSA a less discriminative marker in transition zone-dominant cohorts. The current study included 74% peripheral zone lesions, where PSA is a stronger discriminator.

The study achieved a Radiomics Quality Score (RQS) of 70% (25/36 points) -- above the median for published radiomics studies -- indicating adherence to methodological standards for feature extraction, preprocessing, and validation. Limitations include the relatively small sample size of 214 patients, absence of dynamic contrast-enhanced (DCE) perfusion sequences, and the held-out validation only being applied to the benign-malignant task rather than all grading questions.

TL;DR: DWI b-value selection and scanner field strength are major sources of performance variability across published radiomic studies, explaining why imaging-only AUCs differ widely in the literature even for the same clinical question.
Pages 11-12
Clinical Implications and Path Forward

This study demonstrates that biparametric MRI radiomic features combined with PSA can achieve strong ML classification performance for prostate cancer grading across multiple clinical questions -- including the clinically difficult task of distinguishing ISUP 2 from ISUP 3 cancers within the peripheral zone. The bpMRI approach eliminates the need for contrast injection, reducing scan time, cost, and patient burden compared to full mpMRI, while preserving diagnostic value when PSA is incorporated.

The finding that PSA -- a widely available and routinely measured clinical variable -- dramatically improves all seven ML models has direct clinical implications. It suggests that radiomic models developed without PSA as an input are likely leaving significant discriminative information unused. Future prostate radiomic studies should routinely incorporate PSA and other available clinical variables (age, prostate volume, PSA density) as features alongside imaging-derived metrics.

Among the seven algorithms evaluated, Neural Network and SVM most consistently achieved the highest performance across all classification tasks, with Random Forest and Decision Tree performing strongly on benign-malignant classification. The simpler models -- Logistic Regression, k-NN, and Naive Bayes -- showed moderate but more variable performance, suggesting they are less suited to capturing the complex feature interactions required for fine-grained grading.

The most important recommendation from this study is the need for larger, prospective multicenter datasets with standardized acquisition protocols to validate radiomic ML models before clinical deployment. Standardizing MRI acquisition parameters -- particularly DWI b-values and field strength protocols -- across institutions would allow model results to be compared and combined across studies, and would improve the generalizability of trained models to real-world clinical settings.

TL;DR: Combining biparametric MRI radiomic features with PSA using Neural Network or SVM models achieves strong prostate cancer grading performance, and standardized multicenter validation is the critical next step toward clinical adoption.
Citation: Open Access, . Available at: PMC12387180.