Endometrial cancer is the most common gynecological cancer in developed countries. Early diagnosis is critical - when caught early, survival rates are high, but delayed diagnosis significantly worsens outcomes. In recent years, numerous studies have used AI to improve how endometrial cancer is detected, whether from imaging, histopathology slides, or other data sources.
However, individual studies often reach different conclusions, and it is difficult to know which AI approaches truly work across diverse patient populations. A systematic review and meta-analysis pools results from multiple studies to produce a more reliable overall estimate of how well AI performs for diagnosing endometrial cancer.
The researchers systematically searched major medical databases to find peer-reviewed studies that tested AI systems for diagnosing endometrial cancer. They applied strict inclusion criteria: studies needed a clear reference standard (usually pathology confirmation), report diagnostic accuracy metrics, and use AI as the primary analysis tool. A total of 13 studies met the inclusion criteria, covering more than 6,400 patients and over 438,000 individual samples.
Study quality was assessed using QUADAS-2, a validated tool for evaluating diagnostic accuracy studies. Statistical analyses used random-effects models to account for natural variability between studies. The primary outcomes were pooled sensitivity (ability to detect true cancer cases), specificity (ability to correctly classify non-cancer cases), and the area under the receiver operating characteristic curve (AUC), which combines both measures.
Across all 13 studies, AI demonstrated a pooled sensitivity of 86% and specificity of 92% for diagnosing endometrial cancer. The overall AUC was 0.95, indicating that AI systems correctly distinguish cancer from non-cancer cases about 95% of the time in these studies. These are strong numbers that suggest AI has real diagnostic value.
Breaking results down by AI type revealed an important difference: deep learning methods outperformed traditional machine learning. Deep learning systems achieved a pooled sensitivity of 90%, compared to 81% for conventional machine learning approaches. Specificity was similar between the two groups. This suggests that more sophisticated neural network architectures learn richer diagnostic features from complex medical data.
Despite the impressive numbers, the certainty of the evidence was rated as 'very low' using the GRADE framework, which evaluates how confident we can be in a body of evidence. This low rating reflects several concerns: most included studies had small patient numbers, many did not perform external validation on independent datasets, and there was significant variability in the populations studied and the AI methods used.
High accuracy on a study's own data does not automatically mean the AI will work as well in other hospitals or in routine clinical practice. Many studies also had a high risk of bias in how patients were selected, creating the possibility that results are more optimistic than what would be seen in the real world. These limitations mean the meta-analysis results should be treated as promising early evidence, not as proof of clinical readiness.
The finding that deep learning outperforms traditional machine learning is clinically important. Deep learning models - particularly convolutional neural networks applied to pathology slides or MRI images - can learn complex patterns from raw image data without needing hand-crafted features. This allows them to detect subtle cancer indicators that even experienced clinicians might miss.
Traditional machine learning relies on predefined, hand-selected features that researchers must engineer before training. While these methods can work well, they are constrained by what researchers already know to look for. The superior performance of deep learning suggests that future diagnostic AI tools for endometrial cancer should prioritize neural network architectures over conventional approaches.
The results of this meta-analysis are encouraging but highlight a clear need for higher quality research. Future studies should enroll larger, more diverse patient populations, conduct external validation in independent hospitals, and use consistent imaging and data collection protocols. Without these improvements, it is difficult to determine whether AI systems will work reliably across different healthcare settings.
Reporting standards for AI diagnostic studies also need improvement. Many included studies lacked sufficient detail about their AI methods, training procedures, or validation strategies, making it hard to assess their true reliability. Adopting standardized reporting guidelines - such as STARD-AI - would help future studies be more transparent and comparable, ultimately accelerating the development of trustworthy AI tools for endometrial cancer diagnosis.