Risk prediction of pancreatic cancer in patients with recent onset hyperglycemia: A machine-learning approach

Journal of Clinical Gastroenterology 2023 AI 5 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Page [1, 2]
Using New-Onset Diabetes as an Early Warning Sign

Pancreatic cancer is the third leading cause of cancer death in the United States with a 5-year survival of only 10%. A major reason for this is that most cases are diagnosed at an advanced stage, when treatment options are limited.

New-onset diabetes (NOD) — particularly in people over age 50 — has been identified as a potential early warning sign of underlying pancreatic cancer. When the pancreas develops a tumor, it can disrupt insulin production and cause blood sugar to rise, sometimes years before other symptoms appear.

This study used a machine learning approach applied to electronic health records to build a risk prediction model for pancreatic cancer specifically in patients who recently developed elevated blood sugar (hyperglycemia), aiming to identify who should be screened more urgently.

TL;DR: Machine learning was applied to health records of over 109,000 patients with new-onset hyperglycemia to predict who among them was at highest risk for pancreatic cancer.
Pages 3-3
Analyzing 109,000 Patients with Elevated Blood Sugar

Health plan enrollees aged 50 to 84 who had a first elevated HbA1c (hemoglobin A1c, a measure of blood sugar control) of 6.5% or higher between January 2010 and September 2018 were identified. This resulted in a cohort of 109,266 patients with recent-onset hyperglycemia.

A total of 102 potential predictors were extracted from electronic health records, including demographic data, lab values, medications, and diagnoses. Multiple imputation was used to handle missing data — a common and important challenge in real-world health records.

The random survival forest machine learning algorithm was used to develop and validate risk models, with performance evaluated by the c-index (a measure of discrimination), calibration, sensitivity, specificity, and positive predictive value.

TL;DR: A random survival forest model was trained on 102 variables from electronic health records of 109,266 patients with new elevated blood sugar to predict 3-year pancreatic cancer risk.
Pages 6-6
A Small Variable Set Predicts Cancer Risk Well

The three-year incidence of pancreatic cancer in this hyperglycemic cohort was 1.4 per 1,000 person-years — higher than the general population, confirming that new-onset hyperglycemia enriches for pancreatic cancer risk.

The best-performing models consistently selected age, weight change over the past year, current HbA1c level, and one additional HbA1c measure (from 6, 12, or 18 months prior) as the most predictive variables. C-indexes ranged from 0.81 to 0.82, indicating good discriminative ability.

Among patients in the top 20% of predicted risk, sensitivity was 56-60%, specificity was 80%, and positive predictive value was 2.5-2.6%. This means the model could identify a high-risk group where 1 in 40 patients would develop pancreatic cancer within 3 years.

TL;DR: A four-variable model (age, weight change, and HbA1c values) achieved a c-index of 0.81-0.82, identifying a high-risk group with 2.5% three-year pancreatic cancer incidence.
Page [8, 9]
Targeting Screening at the Moment of Hyperglycemia Diagnosis

The study demonstrates that the point of first elevated HbA1c is a valuable clinical opportunity: patients are already in the healthcare system, a blood test has been performed, and a clinical decision about follow-up is being made. This is an ideal moment to apply a cancer risk algorithm.

A positive predictive value of 2.5% means the model-identified high-risk group has a cancer prevalence 8-10 times higher than the background rate of new-onset hyperglycemia. This level of enrichment may justify targeted screening with endoscopic ultrasound or CT imaging.

The simplicity of the model — requiring only age, weight change, and a few HbA1c values — makes it readily implementable within existing electronic health record systems without specialized equipment or testing.

TL;DR: The four-variable model could be embedded in EHR systems to flag high-risk newly hyperglycemic patients for pancreatic cancer screening at their point of diabetes diagnosis.
Pages 19-19
Machine Learning Makes Early Detection More Actionable

This study demonstrates that a machine learning model using simple, readily available clinical variables can meaningfully stratify pancreatic cancer risk at the point of new-onset hyperglycemia — a clinically important and actionable window.

By identifying a high-risk subgroup among newly hyperglycemic patients, this approach makes targeted pancreatic cancer screening more practical and cost-effective than broad population screening.

Future prospective validation studies are needed, but this research provides a strong foundation for embedding AI risk prediction into routine diabetes management workflows as a cancer early detection strategy.

TL;DR: A simple machine learning model applied at the point of new hyperglycemia diagnosis can identify high-risk patients for pancreatic cancer screening using only routine clinical data.
Citation: Open Access, 2023. Available at: PMC9585151.