The Use of Artificial Intelligence in Urologic Oncology: Current Insights and Challenges

Res Rep Urol 2025 AI 6 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Why Urologic Oncology Needs AI

Urologic cancers represent a significant share of the global cancer burden. Prostate cancer is the fourth most commonly diagnosed cancer worldwide, with approximately 1.4 million new cases annually. Bladder and kidney cancers are also in the global top ten by incidence, and together with testicular and penile cancers they account for roughly 20% of all malignancies in some national populations.

Despite improvements in diagnosis and treatment, urologic oncology faces persistent unresolved problems: overdiagnosis and overtreatment of indolent tumors, high inter-reader variability in imaging and pathology, complex surgical procedures with steep learning curves, and limited tools for personalizing treatment decisions. These challenges cannot be solved by clinical experience alone -- they require pattern recognition across datasets far larger than any individual clinician can internalize.

Artificial intelligence (AI), and specifically machine learning (ML) and deep learning (DL), offers a framework for addressing these problems. AI, ML, and DL can be understood as nested layers: AI is the broad field of machines performing cognitive tasks; ML uses algorithms that learn from data; DL uses large neural networks capable of processing images, video, and complex signals to uncover patterns invisible to human observers.

This narrative review examines the current state of AI across all major genitourinary cancers -- prostate, bladder, renal, testicular, and penile -- as well as AI's emerging roles in robotic surgery, patient communication, and medical writing. It highlights both documented clinical opportunities and the unresolved challenges that must be overcome before these tools can be routinely deployed in patient care.

TL;DR: Urologic cancers affect millions of people worldwide and present persistent diagnostic and treatment challenges that AI and machine learning are uniquely positioned to help address.
Pages 2-5
AI in Prostate Cancer: From MRI to the Operating Room

Prostate cancer is the most explored domain for AI in urology. For detection, the Stanford Prostate Cancer Network (SPCNet) -- a CNN trained on MRI to distinguish aggressive cancer, indolent cancer, and normal tissue -- achieved AUCs of 0.75-0.80 and detected up to 18% of lesions missed by radiologists. A separately validated AI system trained on over 10,000 MRI scans achieved an AUROC of 0.91 compared to 0.86 for radiologists, and detected 6.8% more significant cancers at equivalent specificity in a cohort of 400 cases.

For pathology grading, the PANDA challenge trained algorithms on 10,616 prostate biopsy samples from multiple centers and achieved concordance scores of 0.862 and 0.868 in two independent international validation sets -- performance comparable to expert uropathologists. The FDA-authorized Paige Prostate model, based on multiple instance learning with a recurrent neural network (MIL-RNN), achieved an AUC of 0.99 on its test set and 0.93 on an external validation set of over 12,000 slides.

Beyond detection and grading, AI has shown potential in pathology workflows for tasks including measuring cancer length and volume, quantifying Gleason pattern percentages, identifying perineural invasion, quantifying immunohistochemistry staining, and detecting cribriform growth patterns. Conditional generative adversarial networks (cGANs) have been trained to convert unstained prostate biopsy images into computationally generated H&E-stained images, achieving a Structural Similarity Index (SSIM) of 0.902 -- nearly identical to physically stained slides.

For surgical applications, a deep learning model called DeepSurv was developed to predict urinary continence recovery after robot-assisted radical prostatectomy (RARP) using automated performance metrics (APMs) collected from 100 surgeries. Surgeons classified as high-performing by APM metrics had significantly better patient continence rates at 3 months (47.5% vs 36.7%) and 6 months (68.3% vs 59.2%), demonstrating AI's potential to objectively assess surgical quality and predict patient-specific outcomes.

TL;DR: AI models have achieved radiologist-level or better performance on MRI-based prostate cancer detection, expert uropathologist-level Gleason grading, and can predict surgical outcomes from intraoperative performance data.
Pages 6-8
AI in Bladder and Renal Cancer

For bladder cancer, a CNN-based cystoscopy system trained on 2,102 images achieved 89.7% sensitivity and 94% specificity for distinguishing normal bladder tissue from tumors -- performance that could assist endoscopists in detecting subtle lesions that might otherwise be missed. AI-based analysis of CT and MRI radiomics has been applied to predict tumor staging, and a meta-analysis of 21 imaging-based AI studies showed diagnostic AUCs up to 0.92 for MRI and 0.91 for deep learning models in predicting muscle-invasive disease, though methodological quality was rated poor across most studies.

A multicenter study using AI-powered histologic assays on digitized pathology slides from 944 patients across 12 centers showed that AI could stratify high-risk non-muscle-invasive bladder cancer patients by recurrence risk, progression, BCG-unresponsive disease, and need for cystectomy -- providing prognostic information beyond traditional clinicopathologic factors. High-risk AI classifications were strongly associated with worse outcomes including shorter recurrence-free survival.

For renal cell carcinoma (RCC), a CNN trained on H&E-stained histopathology images achieved 99.1% accuracy for distinguishing normal tissue from RCC, 97.5% for classifying subtypes, and 98.4% for Fuhrman grade prediction. Machine learning models applied to a Korean database of 6,849 patients predicted RCC recurrence with AUCs of 0.836 at 5 years and 0.784 at 10 years post-nephrectomy -- results that could support individualized post-surgical surveillance planning.

A consistent limitation across bladder and renal AI studies is the reliance on small, single-institution retrospective datasets without external validation. A systematic review of AI in renal histopathology emphasized that while performance metrics are impressive, methodological heterogeneity and limited interpretability remain significant barriers. AI in genomics -- such as classifying molecular subtypes of bladder cancer or predicting chemotherapy response -- remains promising but underexplored.

TL;DR: AI shows strong performance for bladder cancer detection and risk stratification, and for renal cancer subtype classification and grading, though most evidence comes from small internal validation studies without external testing.
Pages 9-10
AI in Robotic Surgery and Augmented Reality

Robot-assisted surgery is a cornerstone of urologic oncology treatment, and AI is being applied to improve both performance evaluation and intraoperative guidance. Automated performance metrics (APMs), captured by data-recording devices on robotic systems like the da Vinci, track instrument motion, camera movements, and energy device usage in Cartesian coordinates during live procedures. Machine learning algorithms applied to these APMs can objectively classify surgeons by performance level and provide targeted feedback for training improvement.

A CNN trained on 1,040 intraoperative video frames from three institutions was developed to recognize the seminal vesicle and vas deferens during RARP -- critical anatomical landmarks whose inadvertent injury can cause serious complications. The model demonstrated potential to enhance anatomical recognition, particularly for novice surgeons who may have less experience identifying these structures under robotic optics.

Augmented reality (AR) -- overlaying digital imaging data onto the operative field -- has been integrated into robotic prostatectomy and partial nephrectomy, primarily to assist with tumor margin identification and surgical navigation. AR systems using preoperative CT and MRI models can project three-dimensional anatomical reconstructions onto live tissue views. However, accurately aligning virtual models with moving or deforming soft tissue structures remains a major technical challenge, and current evidence does not clearly demonstrate significant clinical benefit over conventional techniques.

A key limitation across surgical AI applications is the difficulty of real-time tissue tracking: soft tissues like the prostate, kidney, and bladder move and deform during surgery in ways that static preoperative models cannot capture. Solving this problem will require AI systems capable of adapting to dynamic intraoperative anatomy, a challenge that remains partially unsolved and is an active area of research.

TL;DR: AI is beginning to transform robotic surgery by objectively evaluating surgical performance, assisting with intraoperative anatomy recognition, and enabling augmented reality visualization, though real-time tissue tracking remains a major technical challenge.
Pages 11-12
AI for Patient Communication and Medical Writing

Beyond diagnostics and treatment, AI is also transforming how patients receive and understand medical information. Large language models (LLMs) like ChatGPT have been evaluated for answering common prostate cancer questions in plain language. In one study, ChatGPT-generated responses were rated accurate by 71.7-94.3% of urologists, and simplified summaries were rated accurate and sufficient for decision-making in approximately 89% of cases -- suggesting potential as a patient education tool despite not being designed for clinical use.

A purpose-built medical chatbot called PROSCA (Prostate Cancer Communication Assistant) was developed specifically for prostate cancer education with medically validated content and structured dialogues. In an evaluation with nine patients with suspected prostate cancer, 89% reported increased knowledge, 78% found it easy to use without assistance, and all participants expressed interest in reusing the tool in clinical practice. PROSCA differs from general-purpose LLMs in that its content is clinically reviewed and its dialogues are structured for the specific informational needs of prostate cancer patients.

In academic and medical writing, generative AI tools present both opportunities and risks. A survey of urologists found that approximately 53% encountered limitations when using ChatGPT in academic settings, with the most common issues being inaccurate responses (44.7%), lack of specificity (42.4%), and inconsistent outputs (26.5%). Over 50% of respondents acknowledged inaccuracies in AI-generated content, and 62.2% recognized ethical concerns including plagiarism, artificial hallucinations, and difficulties citing AI-generated text appropriately.

A survey of 100 major publishers found that only 24% provided guidance on generative AI use in academic research, and none had developed these guidelines through a structured consensus process -- reflecting a significant regulatory and ethical gap. The medical community broadly recognizes the need for unified policies governing how AI-generated content should be disclosed, validated, and incorporated into clinical documentation and scientific literature.

TL;DR: LLMs show real potential for patient education and communication, but general-purpose tools like ChatGPT lack clinical validation, while purpose-built chatbots like PROSCA demonstrate more dependable performance in urologic patient counseling.
Pages 12-13
Key Limitations and the Path to Clinical Integration

Despite impressive performance metrics across many studies, a consistent pattern of methodological weakness limits the clinical credibility of the current evidence base. The majority of reviewed studies are retrospective, involve single institutions with relatively small patient cohorts, and rely exclusively on internal validation without testing models on independent datasets. High reported accuracies -- some exceeding 99% -- should be interpreted with caution when produced by models that have never been applied to new patients from different institutions or populations.

Algorithmic bias is an underappreciated risk: the majority of published AI models in urologic oncology originate from high-income Western academic centers, raising questions about their relevance and performance in more diverse or resource-limited healthcare settings. Few studies have included adequate representation of non-Western populations, patients with comorbidities, or cases from community hospitals, where imaging protocols and pathology practices may differ substantially from academic centers.

Ethical and regulatory frameworks for AI in medicine remain underdeveloped. Responsibility for AI-generated clinical decisions is poorly defined, privacy and data security for the large patient datasets required to train these models are complex challenges, and regulatory approval pathways -- as navigated by tools like Paige Prostate -- are demanding and slow relative to the pace of technological development. A global survey found that most urologists support establishing clear guidelines, regular auditing of AI systems, and ethical oversight.

The path forward requires four aligned efforts: building prospective multicenter datasets with standardized protocols; conducting rigorous external validation comparing AI tools against established clinical standards; developing explainability mechanisms that make model predictions understandable and actionable for clinicians; and establishing clear ethical and legal frameworks for responsible AI deployment. As AI technology continues to evolve, its greatest value will be realized as a complement to -- not a replacement for -- the clinical judgment of expert urologists and pathologists.

TL;DR: Most urologic AI models lack external validation, reflect geographic bias toward Western academic centers, and operate without adequate ethical or regulatory frameworks -- all of which must be addressed before clinical integration can be broadly trusted.
Citation: Open Access, . Available at: PMC12377376.