Clinical importance of early detection. Endometrial cancer is the fourth most common cancer in women in the United States, after breast, lung, and colorectal cancers. Early diagnosis is critical because adenocarcinoma is more likely to be treated successfully than other female cancers - patients who observe early signs of abnormal vaginal bleeding seek specialist care and receive treatment in earlier stages.
The diagnostic challenge with abnormal uterine bleeding. Abnormal uterine bleeding (AUB) is the most common manifestation of endometrial cancer, especially after menopause. However, it is not a reliable indicator: only 10% of women with AUB have cancer, and 90% of women who undergo aggressive diagnostic procedures do not have cancer. This over-testing burden motivates a predictive screening tool.
Risk factors for endometrial cancer. Established risk factors include high estrogen levels, hypertension, early menarche, late menopause, tamoxifen use, nulliparity, Lynch syndrome, old age (55 or older), obesity, and polycystic ovarian disease. Identifying which women with AUB carry combinations of these risk factors could prioritize biopsy referrals.
Machine learning as a screening tool. Conventional biostatistical methods are not well-suited for managing complex, multivariable clinical data. Machine learning offers flexibility and scalability to identify non-linear risk patterns across diverse variables, enabling the development of a non-invasive predictive model for endometrial cancer from patient characteristics and ultrasound findings.
Study design and cohort. This cross-sectional study enrolled 972 patients with abnormal uterine bleeding referred to a gynecology clinic in Tehran, Iran, between January 2016 and January 2021. Pathological results from dilatation and curettage confirmed 920 benign cases (94.7%) and 52 endometrial cancer cases (5.3%), with mean patient age 45.77 years.
Feature set. Predictors recorded for each patient included age, BMI, type of abnormal uterine bleeding, uterus size on bimanual exam, history of other diseases (diabetes, hypertension, hypothyroidism, PCOS), pregnancy history, menarche age, menopausal age, menopausal status, family history of cancer, tamoxifen use, and endometrial thickness and uterus size on ultrasound.
Four machine learning algorithms. Logistic regression, classification and regression trees (CART), support vector machine (SVM), and artificial neural network (ANN, specifically a multilayer perceptron with backpropagation) were implemented in Python using the Scikit-learn framework. All univariate variables significant at p below 0.05 and frequency above 10% were included in the models.
Train-test methodology. Datasets were split 4:1 into training and test sets. Model performance was evaluated using sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), false positive rate, false negative rate, and overall accuracy on both training and test sets. The final logistic regression model was additionally assessed using the Hosmer-Lemeshow goodness-of-fit test.
Logistic regression performance. Logistic regression achieved the highest sensitivity of 100% on training data and 98% on the test set, with specificity of 98.83% (training) and 98.7% (test). Positive predictive value was 82.97% and 82.85%, and negative predictive value was 100% and 99.89% for training and test sets respectively, making it the top-performing model overall.
Comparison to other models. CART achieved sensitivity of 92.31% (train) and 88.18% (test), with overall accuracy of 99.18% and 98.35%. SVM reached sensitivity of 89.45% (train) and 86.55% (test). ANN achieved sensitivity of 89.92% (train) and 88.55% (test), with the highest specificity at 99.86% (train) and 99.78% (test). Overall accuracy was close across all four models, ranging from 98.35% to 99.33%.
Independent risk factors from logistic regression. Multiple logistic regression identified four independent risk factors: higher BMI (OR=1.34; 95% CI 1.07-1.68), lower menarche age (OR=0.157; 95% CI 0.068-0.363), larger endometrial thickness (OR=1.31; 95% CI 1.16-1.49), and hypertension (OR=15.25; 95% CI 1.02-227.01).
CART decision rules. Nine decision rules were extracted from CART. Menarche age was the most important variable. For patients with menarche before age 11.5, endometrial thickness above 16.5 mm combined with age above 51.1 years was associated with 100% probability of malignancy. For those with menarche after 11.5, hypertension and endometrial thickness drove classification. These rules directly mirror established clinical risk criteria.
Significant univariate differences. Statistically significant differences between benign and malignant groups were found for age, BMI, menarche age, pregnancy history, uterus size on bimanual exam, postmenopausal bleeding, menometrorrhagia, menorrhagia, metrorrhagia, history of cancer, diabetes mellitus, hypertension, PCOS, hypothyroidism, menopausal status, menopause age, and endometrial thickness - all at p below 0.05.
Non-significant variables. Three variables showed no significant group difference: family history of cancer (p=0.259), tamoxifen use (p=0.270), and infertility (p=0.327). This suggests that tamoxifen use and family cancer history, while recognized population-level risk factors, did not differentiate malignant from benign in this specific AUB referral cohort.
Endometrial thickness discrimination. The most striking quantitative difference between groups was endometrial thickness: median 10 mm (IQR 8-13 mm) in benign patients versus 56.5 mm (IQR 19.5-31.75 mm) in malignant cases. This several-fold difference makes endometrial thickness the most powerful single predictor in the dataset.
Postmenopausal bleeding prevalence. Postmenopausal bleeding was present in 82.7% of malignant cases versus only 22.7% of benign cases, confirming its clinical importance as a presenting symptom. However, because the majority of AUB cases overall had benign pathology, postmenopausal bleeding alone remains insufficient for diagnosis without additional predictors.
Why logistic regression outperformed neural networks. This study's finding that logistic regression achieved higher sensitivity than the more complex ANN contrasts with Pergialiotis et al. (2018), who found ANN superior in postmenopausal women. The difference likely reflects cohort composition: this study included both reproductive-age and menopausal women with AUB, while the prior study focused on postmenopausal women alone.
High overall accuracy context. All four models showed overall accuracy above 98%, which is partly explained by the class imbalance (94.7% benign). Sensitivity - the ability to detect true cancer cases - is the clinically more important metric in a cancer screening context, and logistic regression uniquely achieved zero false negatives on the training set.
Clinical applicability. The study proposes using these ML models as a non-invasive first-line screening tool to stratify AUB patients into high-risk and low-risk groups, reducing unnecessary biopsies in low-risk patients while ensuring malignant cases are not missed. This could reduce economic burden and improve resource allocation in gynecology clinics.
Limitations. The study was conducted at a single center with limited access to electronic records from other facilities, restricting dataset size to 972 patients. The sample imbalance (52 EC vs. 920 benign) may affect model generalizability. Multi-center studies with larger datasets are needed to improve sensitivity and validate the models in diverse populations.
Core finding. Among four machine learning models tested on 972 AUB patients, logistic regression achieved the best diagnostic performance with 100% training sensitivity and 98% test sensitivity, 98.83% training specificity, and 99.89% negative predictive value - indicating virtually no missed cancer cases. The four key independent predictors were BMI, early menarche age, endometrial thickness, and hypertension.
Potential to reduce invasive procedures. The high negative predictive value of logistic regression suggests that women scoring low on the model could safely avoid dilation and curettage, reducing procedure-related risks and healthcare costs. CART provides complementary value by generating interpretable decision rules that clinicians can directly apply without computational tools.
Future directions. Multi-center data collection would increase sample size and improve model sensitivity for the minority malignant class. Integration with electronic health records and imaging systems could enable automated real-time risk scoring at point of care, supporting clinical decision-making without requiring dedicated AI expertise from individual physicians.