Artificial Intelligence in Digital Pathology for Bladder Cancer: Hype or Hope? A Systematic Review.

Cancers (Basel) 2023 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-3
The Case for AI in Bladder Cancer Pathology

Bladder cancer diagnosis and prognosis depend heavily on subjective pathological evaluation. Current risk stratification relies on clinicopathological factors like tumor grade and stage, but significant intra- and interobserver variability among pathologists can result in misdiagnosis, incorrect staging, and consequent under- or over-treatment of patients.

Computational pathology offers an objective, reproducible alternative to purely human assessment. By applying machine learning and deep learning algorithms to whole-slide histopathology images, computational pathology can analyze tissue features at scale with consistency unavailable through manual review. One AI-based method for prostate cancer pathology has already received FDA approval, demonstrating clinical viability.

This systematic review analyzed 33 studies selected from 2,285 initial search results. Following PRISMA guidelines, the authors performed a comprehensive search across Embase, Medline, Cochrane, Web of Science, and Google Scholar to map the current state of computational pathology in bladder cancer diagnosis and prognosis prediction.

The review covers both machine learning and deep learning approaches applied to bladder cancer whole-slide images. Machine learning requires hand-crafted feature engineering before training, while deep learning automatically extracts features from raw image data. Both strategies have been applied to bladder cancer histopathology for tissue segmentation, grading, staging, and outcome prediction.

TL;DR: This systematic review synthesizes findings from 33 studies on AI-based computational pathology for bladder cancer, covering tissue segmentation, grading, staging, and prognostic prediction from whole-slide images.
Pages 5-6
Tissue and Cell Segmentation Achievements

AI algorithms can segment bladder tissue types with over 90 percent accuracy. Niazi and colleagues achieved 98, 98, and 94 percent accuracy for lamina propria, muscularis propria, and urothelium segmentation respectively. Wetteland and colleagues reported 96 percent average accuracy segmenting six tissue classes including urothelium, stroma, muscle, blood, damaged tissue, and background.

Multi-magnification analysis improves segmentation accuracy in whole-slide images. Wetteland demonstrated that combining features extracted at different magnification levels outperforms single-resolution analysis. This reflects how pathologists naturally zoom in and out to assess tissue architecture and cellular detail during diagnosis.

Cell nucleus characteristics detectable by AI correlate with cancer prognosis. Changes in nuclear shape, size, polarization, and roundness have been linked to worse clinical outcomes. Glotsos and colleagues developed a nucleus segmentation model achieving 94 percent accuracy, though clinical outcome correlations were not investigated in this early study.

Microvessel segmentation from immunohistochemistry images provides tumor angiogenesis information. Loukas developed a model that segmented microvessels with 87 percent accuracy from CD31-stained images. Intratumoral microvessel density is a marker of metastatic potential, and automated quantification could support objective assessment of this prognostic factor.

TL;DR: AI models achieved over 90 percent accuracy for segmenting bladder tissue types including urothelium, stroma, and muscle, and showed promise for detecting diagnostically relevant cellular and vascular features.
Pages 5-7
Tumor Detection and Grading Performance

Deep learning can discriminate tumor from normal urothelium with AUC up to 0.99. Zhang and colleagues achieved a true positive rate of 0.95 for detecting tumor areas in papillary urothelial carcinoma whole-slide images. Noorbakhsh and colleagues reported AUC of 0.99 for tumor versus normal tissue classification across 19 cancer types, with cross-applicability to bladder cancer.

Models trained on multiple cancer types can classify bladder cancer tissue without disease-specific training data. Jang and colleagues trained models on five other cancer types and achieved 0.94 to 0.98 AUC on bladder cancer whole-slide images, demonstrating morphological similarities that enable transfer of learned features across tumor types.

AI grading models achieved 85 to 96 percent accuracy for high- and low-grade tumor classification. ML-based models using cell nuclei features reached 85 to 96 percent accuracy, while deep learning models achieved 74 to 95 percent. Zhang and colleagues developed an algorithm that outperformed a panel of 17 pathologists, achieving 95 percent grading accuracy versus the pathologists' 84 percent.

Staging models can distinguish Ta from T1 bladder cancer with 96 percent accuracy. Yin and colleagues used nuclear size, cytoplasmic color, nuclear shape, and connective tissue pattern to train a model separating non-invasive from invasive disease. Analysis of the model's features identified desmoplastic reaction as the most important distinguishing characteristic, a finding with potential clinical significance.

Histological growth pattern classification in muscle-invasive disease reached 90 percent accuracy. Garcia and colleagues classified nodular, trabecular, and infiltrative growth patterns from immunohistochemistry images, capturing prognostically relevant architectural features of muscle-invasive bladder cancer that predict recurrence risk.

TL;DR: AI models matched or exceeded expert pathologist performance for bladder cancer grading, and demonstrated the ability to distinguish tumor from normal tissue and classify staging features including desmoplastic reaction at high accuracy.
Pages 8-9
Prognosis Prediction From Histopathology

Nuclear morphometry features predict bladder cancer recurrence with over 90 percent accuracy. Tasoulis and colleagues analyzed maximum area, skewness of area, and maximum concavity of cell nuclei, achieving 92 percent accuracy for recurrence prediction. Tokuyama and colleagues similarly achieved 90 percent accuracy using quantitative nuclear features from pathologist-annotated regions in non-muscle-invasive bladder cancer.

Combining histopathological image features with clinical data significantly improves survival prediction. Chen and colleagues extracted cell nuclei area, contrast, and distribution features from whole-slide images and merged them with clinical information, achieving 81 percent accuracy for 5-year overall survival prediction. This outperformed standard clinicopathological risk stratification systems.

Tumor budding quantification from whole-slide images adds prognostic value beyond clinical parameters. Brieu and colleagues quantified tumor budding in muscle-invasive bladder cancer and showed that combining detected budding features with clinical parameters separated patients into low and high budding groups with significantly different disease-specific survival outcomes.

Cancer-specific survival prediction incorporating immune cell spatial features achieved 89 percent AUC. Gavriel and colleagues used immunofluorescence-stained slides to identify tumor budding, T-cells, macrophages, and PD-L1 expression, then combined spatial and cell density features with clinical information. The resulting model outperformed standard clinicopathological approaches in cancer-specific survival prediction.

Deep learning predicted 5-year recurrence-free survival with AUC of 0.76 in non-muscle-invasive bladder cancer. Lucas and colleagues found that adding image-derived features to clinical variables improved recurrence prediction beyond what clinical data alone could provide, reinforcing the complementary value of computational pathology to existing risk tools.

TL;DR: AI models combining histopathological image features with clinical information predicted bladder cancer recurrence, overall survival, and cancer-specific survival with accuracies typically ranging from 80 to 92 percent, outperforming standard clinicopathological models.
Page 9
Biomarker and Molecular Subtype Detection

Deep learning can predict molecular subtypes of muscle-invasive bladder cancer from standard H&E slides. Woerl and colleagues achieved 75 percent accuracy predicting luminal, basal, neuronal, and stroma-rich molecular subtypes without molecular testing. Crucially, when pathologists were shown the model's highlighted regions, their own molecular subtype prediction accuracy improved from 38 to 59 percent.

Specific morphological features correlate with molecular subtypes and can be identified by AI. The model highlighted hyperchromatic nuclei with low pleomorphism for double-negative subtypes, large pleomorphic nuclei with atypical nucleoli for basal tumors, papillary structures for luminal tumors, and small infiltrating cell nests for luminal p53-like tumors. These associations represent novel pathological insights enabled by AI analysis.

FGFR mutation status can be predicted from routine histology images. Velmahos and colleagues predicted FGFR-activating mutations with 0.82 sensitivity by estimating the tumor-infiltrating lymphocyte proportion as a surrogate. Loeffler and colleagues directly detected FGFR3 mutations from H&E images with 0.72 AUC, offering a rapid pre-screening tool before expensive molecular testing.

Tumor-infiltrating lymphocyte mapping from whole-slide images achieved 0.95 AUC across 13 cancer types. Saltz and colleagues developed a pan-cancer TIL detection model using pathologist-labeled training data, with applicability to bladder cancer. Since TIL density correlates with immune checkpoint therapy response, automated TIL quantification could guide immunotherapy selection.

TL;DR: AI detected bladder cancer molecular subtypes from standard pathology slides, predicted FGFR mutations, and mapped tumor-infiltrating lymphocytes, offering surrogate molecular profiling without the cost and complexity of genomic testing.
Pages 14-16
Gaps Between Research and Clinical Practice

Most studies analyzed manually selected regions of interest rather than entire whole-slide images. Focusing on pre-selected regions risks missing diagnostically relevant information elsewhere on the slide. It also produces models that depend on specialist annotation, reducing scalability and increasing the risk of selection bias that may not reflect real clinical variability.

The Cancer Genome Atlas was used in 42 percent of included studies, raising concerns about institutional bias. Models trained primarily on TCGA data may learn institution-specific tissue preparation and staining patterns as spurious predictive features. These artifacts can artificially inflate performance metrics on internal validation sets while failing to generalize to external clinical datasets.

The lack of outcome interpretability in deep learning models limits clinical trust and adoption. Clinicians cannot make healthcare decisions based on unexplained algorithmic outputs. Explainability tools that highlight which image features drove a prediction are essential for building clinician confidence and enabling regulatory approval for diagnostic AI in clinical use.

Dataset sizes were frequently small, with many studies analyzing fewer than 200 patients. Small cohorts increase overfitting risk and reduce statistical power to detect true predictive signals. The review recommends datasets of more than 100 patients as a minimum, with prospective follow-up long enough to capture recurrence, progression, and survival events.

TL;DR: The most critical gaps in bladder cancer computational pathology are overreliance on small and biased datasets, limited whole-slide analysis, absence of external validation, and insufficient model interpretability for clinical deployment.
Pages 16-17
Recommendations for Clinical Integration

A standardized framework for data collection, processing, and reporting is urgently needed. The review's Table 2 outlines recommendations spanning data collection through clinical implementation. Key priorities include publishing datasets and algorithms openly, reporting comprehensive performance metrics including AUC, F1 score, sensitivity and specificity, and tracking every step of data preprocessing transparently.

Models should be validated on demographically diverse external patient cohorts. Cross-demographic validation ensures that models generalize across different patient populations, staining protocols, and scanner types. Models that only perform well on their development dataset provide limited real-world clinical value.

Explainable AI methods should be integrated into all diagnostic and prognostic models. Feature attribution and attention visualization tools allow clinicians to understand why an AI system reached a particular conclusion. This transparency is legally and ethically necessary and enables detection of spurious feature learning from artifacts or institutional patterns.

Fusion models combining histopathological images with clinical and molecular data represent the most promising direction. Multiple included studies demonstrated that combining image-derived features with clinical variables consistently improved prediction accuracy over either source alone. Integrating clinicopathological, histopathological, and molecular subtyping data could open new horizons for precision medicine in bladder cancer.

Adaptive model monitoring and updating after clinical deployment is essential for sustained accuracy. Patient populations, staining protocols, and scanner technologies evolve over time. Continuous monitoring ensures models remain accurate as clinical contexts change, and updating with new data prevents performance degradation in real-world settings.

TL;DR: The review provides actionable recommendations for integrating computational pathology into bladder cancer clinical practice, centered on standardized data practices, external validation, explainability, and multimodal data fusion.
Page 17
AI as Hope, Not Hype, for Bladder Cancer

Computational pathology has demonstrated real potential to improve bladder cancer diagnosis and prognostic prediction. Across 33 studies, AI models achieved high accuracy for tissue segmentation, tumor detection, grading, staging, molecular subtype identification, and clinical outcome prediction. Several models outperformed pathologists or standard clinicopathological tools on their respective tasks.

AI is not designed to replace pathologists but to augment their expertise. The goal of computational pathology is to provide objective, reproducible second opinions that guide pathologist attention to diagnostically important regions, reduce observer variability, and identify novel prognostic markers that human visual inspection alone cannot reliably detect.

Addressing the black box challenge is essential for AI adoption in clinical pathology. Regulatory acceptance, legal accountability, and clinician trust all require that AI decisions be interpretable and justifiable. Explainable AI techniques that translate model outputs into clinically relevant findings are a prerequisite for moving from research to clinical implementation.

When computational pathology meets bladder cancer, hope prevails over hype. The field has moved beyond proof-of-concept demonstrations to produce models with measurable clinical accuracy improvements. With larger datasets, external validation, standardized practices, and interpretable designs, computational pathology holds genuine promise to transform bladder cancer management and support precision medicine.

TL;DR: The systematic review concludes that AI-based computational pathology represents genuine hope for improving bladder cancer diagnosis and prognosis, provided that key challenges around data standardization, model interpretability, and external validation are addressed.
Citation: Open Access, 2023. Available at: PMC10526515.