Artificial intelligence (AI) tools for renal cell carcinoma (RCC) detection have proliferated rapidly over the past decade, driven by advances in deep learning and the increasing availability of large annotated imaging datasets. However, the clinical reliability of these tools has been difficult to assess because individual studies vary widely in methodology, patient population, imaging modality, and outcome definition.
This systematic review and meta-analysis synthesized evidence from 64 published studies examining AI-based RCC detection, with 31 studies contributing quantitative data to the meta-analysis. By pooling results across studies, the authors generated summary estimates of AI performance that are more reliable than any single-study finding and allow direct comparison between AI and clinician performance.
The review addressed three primary questions: how well AI performs on internal validation (testing on data from the same institution and period as training), how well AI performs on external validation (testing on data from different institutions or time periods), and how AI performance compares to clinician performance on the same task.
A comprehensive literature search was conducted across major medical and computer science databases to identify studies reporting AI performance for RCC detection tasks including tumor identification, localization, segmentation, or classification. Studies were included if they reported sensitivity and specificity or sufficient data to calculate these metrics.
The 64 included studies were assessed for methodological quality using established frameworks for AI diagnostic accuracy studies. Key quality dimensions included study design (prospective vs. retrospective), sample size, reporting of class distribution, validation strategy (internal vs. external), and whether AI was compared directly to clinician performance.
For the meta-analysis, the authors used a bivariate random-effects model, which accounts for the correlation between sensitivity and specificity and the heterogeneity between studies. This is the recommended approach for meta-analysis of diagnostic test performance and produces summary receiver operating characteristic (sROC) curves alongside pooled sensitivity and specificity estimates with confidence intervals.
Across studies reporting internal validation results, the pooled sensitivity of AI for RCC detection was 85% (95% confidence interval 82% to 87%) and the pooled specificity was 76% (95% CI 70% to 80%). This level of performance represents strong but imperfect discrimination, with approximately 1 in 7 malignant lesions missed and nearly 1 in 4 benign lesions incorrectly flagged.
Substantial heterogeneity was observed among studies, reflecting differences in the AI architectures used, the specific RCC detection tasks evaluated (e.g., detection vs. segmentation vs. subtype classification), and patient population characteristics. High heterogeneity is common in meta-analyses of AI diagnostic studies and reflects the diversity of real-world application contexts.
Internal validation estimates tend to be optimistic because the model has been exposed to data from the same distribution as the test set. The gap between internal and external validation performance is a key indicator of how much overfitting and distribution shift affect real-world performance.
Among studies reporting external validation results, pooled sensitivity was 80% (95% CI 73% to 84%) and pooled specificity was 90% (95% CI 84% to 93%). Notably, specificity was substantially higher in external validation than in internal validation, while sensitivity dropped modestly.
This pattern suggests that AI models trained on internal data may be calibrated to maximize sensitivity at the expense of specificity, or that external validation datasets tend to include more clearly negative cases that are easier for the model to correctly classify as benign. Alternatively, publication bias may contribute if authors preferentially report external validation results when they show high specificity.
The maintenance of reasonable sensitivity at 80% in external validation is encouraging, as sensitivity is the more clinically critical metric for a cancer detection task where missed malignancies carry serious consequences. However, the confidence intervals for both metrics are wide, reflecting the limited number of studies contributing to the external validation meta-analysis.
A subset of studies compared AI performance directly to clinician performance on the same patient datasets, enabling a controlled comparison. Clinician pooled sensitivity was 79% (95% CI 72% to 85%) and pooled specificity was 60% (95% CI 49% to 70%).
Comparing these estimates to AI performance suggests that AI achieves similar or slightly better sensitivity than clinicians while substantially outperforming clinicians in specificity. The clinician specificity of only 60% indicates that experienced physicians frequently label benign lesions as potentially malignant, a tendency that drives unnecessary biopsies and surgeries. AI specificity in the 76-90% range represents a meaningful improvement.
However, it is important to note that clinicians in these studies typically operated without the assistance of structured quantitative analysis tools, and the comparison may not reflect what would happen if AI were used as a decision support adjunct rather than a replacement. Human-AI collaboration models may outperform either alone.
Despite the generally encouraging performance metrics, the review identified several important gaps in the current evidence base. Most studies were retrospective and single-institution, limiting their applicability to broader clinical settings. Prospective multi-center validation trials are needed to assess real-world performance and operational requirements before clinical deployment.
Reporting standards for AI diagnostic studies were often inadequate. Many studies did not report confidence intervals, did not describe class imbalance in their datasets, or did not use appropriate methods for handling imbalanced data. These reporting deficiencies make it difficult to assess the reliability of individual study findings and contribute to the heterogeneity observed in the meta-analysis.
The review concludes that while AI shows genuine promise for RCC detection and may particularly help reduce false positive rates compared to current clinical practice, the field needs larger prospective studies with standardized protocols, multi-center validation, and direct integration with clinical workflows before AI tools can be confidently recommended for routine use in RCC detection.