Systematic Review and Meta-Analysis of Artificial Intelligence for Renal Cell Carcinoma Detection

BMC Cancer 2025 AI 6 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Scope and Purpose of the Systematic Review

Artificial intelligence (AI) tools for renal cell carcinoma (RCC) detection have proliferated rapidly over the past decade, driven by advances in deep learning and the increasing availability of large annotated imaging datasets. However, the clinical reliability of these tools has been difficult to assess because individual studies vary widely in methodology, patient population, imaging modality, and outcome definition.

This systematic review and meta-analysis synthesized evidence from 64 published studies examining AI-based RCC detection, with 31 studies contributing quantitative data to the meta-analysis. By pooling results across studies, the authors generated summary estimates of AI performance that are more reliable than any single-study finding and allow direct comparison between AI and clinician performance.

The review addressed three primary questions: how well AI performs on internal validation (testing on data from the same institution and period as training), how well AI performs on external validation (testing on data from different institutions or time periods), and how AI performance compares to clinician performance on the same task.

TL;DR: This systematic review and meta-analysis of 64 studies pooled AI performance for RCC detection across internal validation, external validation, and head-to-head comparison with clinicians.
Pages 2-4
Study Selection and Quality Assessment

A comprehensive literature search was conducted across major medical and computer science databases to identify studies reporting AI performance for RCC detection tasks including tumor identification, localization, segmentation, or classification. Studies were included if they reported sensitivity and specificity or sufficient data to calculate these metrics.

The 64 included studies were assessed for methodological quality using established frameworks for AI diagnostic accuracy studies. Key quality dimensions included study design (prospective vs. retrospective), sample size, reporting of class distribution, validation strategy (internal vs. external), and whether AI was compared directly to clinician performance.

For the meta-analysis, the authors used a bivariate random-effects model, which accounts for the correlation between sensitivity and specificity and the heterogeneity between studies. This is the recommended approach for meta-analysis of diagnostic test performance and produces summary receiver operating characteristic (sROC) curves alongside pooled sensitivity and specificity estimates with confidence intervals.

TL;DR: A bivariate random-effects model was applied to pool sensitivity and specificity from 31 quantitative studies, with separate analyses for internal validation, external validation, and clinician comparison.
Pages 5-7
Internal Validation Performance: 85% Sensitivity, 76% Specificity

Across studies reporting internal validation results, the pooled sensitivity of AI for RCC detection was 85% (95% confidence interval 82% to 87%) and the pooled specificity was 76% (95% CI 70% to 80%). This level of performance represents strong but imperfect discrimination, with approximately 1 in 7 malignant lesions missed and nearly 1 in 4 benign lesions incorrectly flagged.

Substantial heterogeneity was observed among studies, reflecting differences in the AI architectures used, the specific RCC detection tasks evaluated (e.g., detection vs. segmentation vs. subtype classification), and patient population characteristics. High heterogeneity is common in meta-analyses of AI diagnostic studies and reflects the diversity of real-world application contexts.

Internal validation estimates tend to be optimistic because the model has been exposed to data from the same distribution as the test set. The gap between internal and external validation performance is a key indicator of how much overfitting and distribution shift affect real-world performance.

TL;DR: Pooled internal validation performance was 85% sensitivity (CI 82-87) and 76% specificity (CI 70-80), with substantial heterogeneity across studies.
Pages 7-9
External Validation: Specificity Rises While Sensitivity Drops

Among studies reporting external validation results, pooled sensitivity was 80% (95% CI 73% to 84%) and pooled specificity was 90% (95% CI 84% to 93%). Notably, specificity was substantially higher in external validation than in internal validation, while sensitivity dropped modestly.

This pattern suggests that AI models trained on internal data may be calibrated to maximize sensitivity at the expense of specificity, or that external validation datasets tend to include more clearly negative cases that are easier for the model to correctly classify as benign. Alternatively, publication bias may contribute if authors preferentially report external validation results when they show high specificity.

The maintenance of reasonable sensitivity at 80% in external validation is encouraging, as sensitivity is the more clinically critical metric for a cancer detection task where missed malignancies carry serious consequences. However, the confidence intervals for both metrics are wide, reflecting the limited number of studies contributing to the external validation meta-analysis.

TL;DR: External validation showed pooled sensitivity of 80% (CI 73-84) and specificity of 90% (CI 84-93), with notably higher specificity than seen in internal validation.
Pages 9-11
AI vs. Clinician Performance Comparison

A subset of studies compared AI performance directly to clinician performance on the same patient datasets, enabling a controlled comparison. Clinician pooled sensitivity was 79% (95% CI 72% to 85%) and pooled specificity was 60% (95% CI 49% to 70%).

Comparing these estimates to AI performance suggests that AI achieves similar or slightly better sensitivity than clinicians while substantially outperforming clinicians in specificity. The clinician specificity of only 60% indicates that experienced physicians frequently label benign lesions as potentially malignant, a tendency that drives unnecessary biopsies and surgeries. AI specificity in the 76-90% range represents a meaningful improvement.

However, it is important to note that clinicians in these studies typically operated without the assistance of structured quantitative analysis tools, and the comparison may not reflect what would happen if AI were used as a decision support adjunct rather than a replacement. Human-AI collaboration models may outperform either alone.

TL;DR: Clinicians achieved pooled sensitivity of 79% and specificity of 60%, suggesting AI offers similar sensitivity but substantially better specificity than unassisted clinical judgment.
Pages 12-14
Evidence Gaps and Requirements for Clinical Translation

Despite the generally encouraging performance metrics, the review identified several important gaps in the current evidence base. Most studies were retrospective and single-institution, limiting their applicability to broader clinical settings. Prospective multi-center validation trials are needed to assess real-world performance and operational requirements before clinical deployment.

Reporting standards for AI diagnostic studies were often inadequate. Many studies did not report confidence intervals, did not describe class imbalance in their datasets, or did not use appropriate methods for handling imbalanced data. These reporting deficiencies make it difficult to assess the reliability of individual study findings and contribute to the heterogeneity observed in the meta-analysis.

The review concludes that while AI shows genuine promise for RCC detection and may particularly help reduce false positive rates compared to current clinical practice, the field needs larger prospective studies with standardized protocols, multi-center validation, and direct integration with clinical workflows before AI tools can be confidently recommended for routine use in RCC detection.

TL;DR: Most studies are retrospective and single-institution; prospective multi-center trials with standardized reporting are needed before AI tools for RCC detection can be recommended for clinical use.
Citation: Open Access, 2025. Available at: PMC11773916.