Pancreatic cancer is the third leading cause of cancer death in the United States with a 5-year survival of only 10%. A major reason for this is that most cases are diagnosed at an advanced stage, when treatment options are limited.
New-onset diabetes (NOD) — particularly in people over age 50 — has been identified as a potential early warning sign of underlying pancreatic cancer. When the pancreas develops a tumor, it can disrupt insulin production and cause blood sugar to rise, sometimes years before other symptoms appear.
This study used a machine learning approach applied to electronic health records to build a risk prediction model for pancreatic cancer specifically in patients who recently developed elevated blood sugar (hyperglycemia), aiming to identify who should be screened more urgently.
Health plan enrollees aged 50 to 84 who had a first elevated HbA1c (hemoglobin A1c, a measure of blood sugar control) of 6.5% or higher between January 2010 and September 2018 were identified. This resulted in a cohort of 109,266 patients with recent-onset hyperglycemia.
A total of 102 potential predictors were extracted from electronic health records, including demographic data, lab values, medications, and diagnoses. Multiple imputation was used to handle missing data — a common and important challenge in real-world health records.
The random survival forest machine learning algorithm was used to develop and validate risk models, with performance evaluated by the c-index (a measure of discrimination), calibration, sensitivity, specificity, and positive predictive value.
The three-year incidence of pancreatic cancer in this hyperglycemic cohort was 1.4 per 1,000 person-years — higher than the general population, confirming that new-onset hyperglycemia enriches for pancreatic cancer risk.
The best-performing models consistently selected age, weight change over the past year, current HbA1c level, and one additional HbA1c measure (from 6, 12, or 18 months prior) as the most predictive variables. C-indexes ranged from 0.81 to 0.82, indicating good discriminative ability.
Among patients in the top 20% of predicted risk, sensitivity was 56-60%, specificity was 80%, and positive predictive value was 2.5-2.6%. This means the model could identify a high-risk group where 1 in 40 patients would develop pancreatic cancer within 3 years.
The study demonstrates that the point of first elevated HbA1c is a valuable clinical opportunity: patients are already in the healthcare system, a blood test has been performed, and a clinical decision about follow-up is being made. This is an ideal moment to apply a cancer risk algorithm.
A positive predictive value of 2.5% means the model-identified high-risk group has a cancer prevalence 8-10 times higher than the background rate of new-onset hyperglycemia. This level of enrichment may justify targeted screening with endoscopic ultrasound or CT imaging.
The simplicity of the model — requiring only age, weight change, and a few HbA1c values — makes it readily implementable within existing electronic health record systems without specialized equipment or testing.
This study demonstrates that a machine learning model using simple, readily available clinical variables can meaningfully stratify pancreatic cancer risk at the point of new-onset hyperglycemia — a clinically important and actionable window.
By identifying a high-risk subgroup among newly hyperglycemic patients, this approach makes targeted pancreatic cancer screening more practical and cost-effective than broad population screening.
Future prospective validation studies are needed, but this research provides a strong foundation for embedding AI risk prediction into routine diabetes management workflows as a cancer early detection strategy.