Artificial Intelligence Algorithms and Their Current Role in the Identification and Comparison of Gleason Patterns in Prostate Cancer Histopathology: A Comprehensive Review

Diagnostics (Basel) 2024 Digital Pathology 7 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Page 2
The Gleason Grading System: A Critical but Imperfect Tool

The Gleason grading system, developed in 1966 by Donald Gleason, is the standard method for assessing how aggressive a prostate cancer is. It scores tumors based on the architectural pattern of cancer cells -- specifically how disorganized and different they appear from normal prostate tissue.

A Gleason score is calculated by adding the two most prevalent tumor patterns, each rated on a scale of 1 to 5. For example, a score of 3+4=7 means the dominant pattern is grade 3, while grade 4 is the next most common. Even with the same total score, a 4+3=7 cancer is more aggressive than a 3+4=7 because the higher-grade pattern dominates.

In 2014, a major revision introduced five Grade Groups to replace the older Gleason score ranges: Grade Group 1 (Gleason 6 and below) through Grade Group 5 (Gleason 9-10). This system provides a cleaner, more clinically meaningful way to communicate cancer severity and guide treatment decisions.

Despite the system's widespread use, a significant problem persists: even experienced pathologists frequently disagree on Gleason grades for the same biopsy. This inter-observer variability can lead to inconsistent treatment recommendations and remains one of the most critical unresolved problems in prostate cancer diagnostics.

TL;DR: The Gleason grading system is the cornerstone of prostate cancer diagnosis, but persistent disagreement between pathologists on the same biopsies creates a pressing need for more consistent and objective tools.
Pages 2-3
How Artificial Intelligence Is Entering the Pathology Lab

Artificial intelligence (AI), particularly deep learning powered by convolutional neural networks (CNNs), has emerged as a powerful approach to analyzing microscopy images with a level of consistency and speed that no human pathologist can match. These systems learn to recognize visual patterns from thousands of labeled training images.

The rise of digital pathology -- where glass slides are converted into high-resolution digital images called whole slide images (WSIs) -- has made it possible to apply AI directly to the same data that pathologists view. AI tools can scan an entire biopsy slide in seconds and provide systematic, reproducible assessments.

This review examined the published literature on AI for Gleason pattern identification, drawing from studies found in PubMed and Google Scholar. The goal was to assess and compare the outcomes and reliability of multiple AI approaches, from simple machine learning models to advanced deep neural networks and multi-omics systems.

TL;DR: Digital pathology has enabled AI systems to analyze prostate biopsy images at scale, offering a potential solution to the inconsistency problem that has long affected Gleason grading.
Pages 3-5
Machine Learning Models for Automated Gleason Classification

Several machine learning approaches have demonstrated strong performance for Gleason classification. A deep learning model built on stimulated Raman scattering (SRS) microscopy -- which generates tissue images without the need for traditional staining -- achieved 85.7% accuracy in classifying Gleason patterns from 61 patients, with 84.4% accuracy on 22 independent validation cases.

Studies comparing ensemble models with deep residual networks found that both approaches achieved approximately 88-89% accuracy in distinguishing cancerous from non-cancerous tissue. Importantly, only deep learning models (ResNets) could reliably differentiate specific Gleason patterns, while traditional quantitative models struggled with this finer distinction.

A deep learning pipeline using digitized biopsy specimens showed a classification accuracy of approximately 80% for assigning Grade Groups to prostate biopsies, with precision and negative predictive values near 94%. The system's agreement with pathologists matched the level of agreement between pathologists themselves -- a meaningful benchmark.

The CorrSigNIA system, which combines MRI images with whole-mount histopathology, achieved over 80% accuracy in cancer identification (AUC 0.81) and performed even better at identifying clinically significant cancers (AUC 0.82-0.86), demonstrating that fusing imaging modalities can improve diagnostic precision beyond either alone.

TL;DR: Early AI models using CNNs and ensemble approaches achieved 80-89% accuracy in Gleason classification, with some fusing imaging modalities to better distinguish aggressive from indolent cancers.
Pages 4-6
Advanced Neural Networks: DeepDx, U-Net, and Beyond

The DeepDx Prostate AI algorithm, trained on 1,133 prostate core needle biopsies, achieved a Cohen's kappa score of 0.91 -- indicating near-perfect agreement with the expert reference standard. For cancer detection, DeepDx achieved 96% accuracy, 99.7% sensitivity, and 88% specificity, making it one of the strongest performers reviewed.

In a study where pathologists used DeepDx as an assistive tool, agreement on Gleason patterns 4 and 5 improved substantially -- the kappa value rose from 0.741 to 0.925 for individual patterns. Diagnostic time was also reduced by 33.9%, demonstrating AI's ability to make pathologists both more accurate and more efficient simultaneously.

A modified U-Net architecture designed for prostate cancer segmentation agreed with pathologist diagnoses on 21 out of 22 sections and achieved an AUC of 96.8% in distinguishing cancerous from non-cancerous areas. Pixel-level accuracy ranged from 93-97% depending on the type of analysis, showing the model's ability to identify precisely where cancer is located within a slide.

A large-scale deep learning system by Strom et al., trained on over 6,600 digitized biopsy slides, achieved an AUC of 0.997 for cancer detection on the test dataset and 0.986 on an external validation dataset -- among the highest reported values in the field -- and correlated with pathologists' tumor length measurements at 0.96, matching expert-level anatomical precision.

TL;DR: Advanced neural networks including DeepDx and U-Net achieved near-expert-level accuracy in cancer detection and Gleason grading, while also making pathologists more accurate and faster when used as assistive tools.
Pages 6-8
Multi-Omics and Radiomic Approaches for Gleason Prediction

Beyond image analysis alone, researchers are integrating multiple biological data types to predict Gleason scores. A multi-omics CNN model incorporating copy number alterations, DNA methylation, and gene expression data achieved a remarkable 98.89% accuracy and an AUC of 0.9996 -- surpassing all prior models reviewed. This highlights that combining genomic and epigenomic data can provide a more complete picture of cancer aggressiveness than imaging alone.

A multi-omics study of 146 prostate cancer patients combined radiomics (from PET/MRI scans), genomics (whole exome sequencing), and proteomics (immunohistochemistry). The best-performing algorithm improved sensitivity by 13%, negative predictive value by 7%, and AUC by 4% compared to standard needle biopsy-based Gleason grading.

Radiomic features extracted from MRI using novel mathematical approaches such as the joint intensity matrix (JIM) achieved AUC values of 78-82% for predicting different Gleason score ranges, complementing standard gray-level image analysis. This suggests that hidden texture patterns in MRI images encode meaningful information about tumor biology.

Machine learning applied to multi-parametric ultrasound -- combining B-mode imaging, shear-wave elastography, and dynamic contrast-enhanced ultrasound -- achieved an AUC of 0.75 for detecting prostate cancer and 0.90 for detecting clinically significant cancer, showing that AI can extract useful diagnostic information from ultrasound, a more widely available imaging modality than MRI.

TL;DR: Combining genomic, radiomic, and imaging data in multi-omics AI models achieves the highest reported accuracy for Gleason prediction, outperforming image-based analysis alone.
Pages 11-13
Practical Barriers to Clinical Adoption

Despite impressive performance in research settings, most AI systems face significant barriers before they can be routinely used in clinical practice. Dataset limitations are the most commonly cited problem: models trained on small or demographically narrow datasets may not generalize well to different patient populations, tissue preparation techniques, or scanner hardware.

Overfitting is a recurring concern -- models may learn to recognize the specific patterns in their training data very well, but fail to perform accurately when presented with new slides from different hospitals or patient groups. This has been documented in multiple studies reviewed, including ensemble U-Net models that excelled on training data but struggled with unseen cases.

Standardization is lacking across the field. Input image dimensions, staining protocols, scanner models, and annotation methods vary from study to study. A model trained on small image patches fed into ResNet, for example, may generate features that do not represent actual tissue architecture well. The absence of common technical standards makes comparing and combining results across studies unreliable.

Some commercial systems, including Paige Prostate (FDA-approved) and AIRA Prostate, have demonstrated high real-world performance, but commercial algorithms were found to underestimate certain cases more often than academic models -- likely because commercial algorithms are optimized for clinical decision-making rather than kappa coefficients used in research benchmarks.

TL;DR: Small and biased training datasets, overfitting, lack of standardization, and limited real-world validation remain the primary obstacles to widespread clinical adoption of AI-based Gleason grading.
Page 13
AI as a Complementary Tool: The Path Forward

The current consensus in the field is that AI should be treated as a complementary tool rather than a replacement for experienced pathologists. The value of AI lies in providing consistent second opinions, reducing diagnostic time, flagging difficult cases, and helping less experienced pathologists achieve expert-level accuracy.

For AI to fulfill its potential in clinical prostate cancer diagnostics, several priorities must be addressed: developing larger and more demographically diverse training datasets, improving model interpretability so pathologists understand why the AI is making a particular decision, and conducting prospective multicenter validation studies that test performance across different hospitals and patient populations.

The combination of AI with multi-omics data represents the most ambitious and promising frontier. Integrating imaging, genomics, and proteomics into unified models could eventually enable fully personalized risk assessment that goes far beyond what any current Gleason grading can provide.

As training datasets continue to grow and algorithms become more refined, AI systems are expected to play an increasingly central role in the era of precision medicine for prostate cancer -- not by replacing human judgment, but by augmenting it with consistent, scalable, and data-rich analysis capabilities.

TL;DR: AI is a powerful and rapidly maturing assistant for Gleason grading that improves accuracy and efficiency, but requires larger diverse datasets, standardization, and multicenter validation before replacing human pathologist judgment.
Citation: Open Access, . Available at: PMC11475684.