Machine Learning Models for Pancreatic Cancer Risk Prediction Using Electronic Health Record Data - A Systematic Review and Assessment

American Journal of Gastroenterology 2024 AI 5 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Page [1, 2]
Surveying the Landscape of AI-Based Pancreatic Cancer Risk Tools

Pancreatic cancer is projected to become the second leading cause of cancer death in the US by 2030. Despite this, fewer than 20% of patients are eligible for curative surgery at diagnosis because most cases are caught too late. Identifying high-risk individuals earlier — and getting them into screening programs — could significantly change these odds.

Electronic Health Records (EHRs) contain a wealth of patient information: lab results, diagnoses, medications, notes, and more. Machine learning can analyze this data to predict who is at risk years before symptoms appear. But how well do these models actually work?

Mayo Clinic researchers conducted a systematic review of 30 published studies covering 169,149 pancreatic cancer cases to critically assess the state of AI-based risk prediction from EHR data.

TL;DR: A systematic review of 30 studies examined how well machine learning models predict pancreatic cancer risk using electronic health record data.
Page [2, 3]
How the Review Was Conducted

The research team searched six major medical databases for studies published between January 2012 and February 2024, looking for AI or ML models trained on EHR data to predict pancreatic cancer. After screening, 30 studies met the inclusion criteria.

Data extraction followed the CHARMS checklist (CHecklist for critical Appraisal and data extraction for systematic Reviews of prediction Modelling Studies). Risk of bias and applicability were evaluated using the PROBAST tool. Models were categorized into three groups: linear models (Group A), non-linear/ensemble models (Group B), and deep learning models (Group C).

The C-index — equivalent to AUROC for survival data — was used as the primary performance metric. Because studies varied enormously in their design, the team standardized comparisons by using the smallest exclusion window and shortest prediction time window for each study.

TL;DR: The review used standardized checklists to evaluate 30 AI risk prediction studies covering over 169,000 pancreatic cancer patients.
Pages 4-4
Model Performance and Key Findings Across 30 Studies

Machine learning model discrimination performance (C-index) ranged widely from 0.57 to 1.0 across the 30 studies. Logistic regression was the most commonly used method (18 of 30 studies). Most models (20 of 30) relied on curated sets of known risk factors selected by clinical experts, such as new-onset diabetes, weight loss, and abnormal lab values.

Models using curated risk factors and those using non-curated EHR features performed similarly on average (mean C-index of 0.81 vs. 0.80). This suggests that carefully chosen clinical predictors are as informative as broader data-mining approaches, at least with current methods.

Missing data handling was poorly reported: only 16 of 30 studies described how they dealt with missing values. Very few studies implemented explainable AI techniques, making it difficult to understand which factors were driving model predictions. And nearly no studies reported whether they excluded data from the period immediately before diagnosis — a critical issue since many early cancer symptoms can look like the disease itself.

TL;DR: AI models for pancreatic cancer risk showed C-indices from 0.57 to 1.0, with logistic regression dominant, but most studies had significant methodological gaps.
Page [5, 6]
Gaps That Need to Be Fixed Before These Models Can Be Used Clinically

The most common methodological weakness was failure to report an 'exclusion time interval' — the practice of removing data from the months immediately before cancer diagnosis. Without this, models may appear more accurate because they're picking up on early cancer symptoms rather than true long-term risk factors.

Few studies used explainable AI techniques (such as SHAP values) to identify which specific EHR variables drove predictions. This opacity is problematic for clinical adoption, since physicians need to understand why a model flags a patient as high-risk.

Deep learning models (Group C) were less likely to use curated risk factors and more likely to use raw EHR data patterns — suggesting they may be discovering novel risk signals. However, they were also the most opaque models, and their clinical interpretability remains a challenge.

TL;DR: Key gaps include poor reporting of missing data, lack of explainability, and failure to account for the time immediately before cancer diagnosis.
Page [6, 7]
The Path Forward: Better AI Models for Pancreatic Cancer Screening

AI/ML models built on curated clinical risk factors perform reasonably well and may be ready for near-term use in identifying high-risk cohorts for targeted pancreatic cancer screening, provided they are validated on real-world primary care datasets.

The most exciting frontier is combining structured EHR data (lab values, diagnoses) with unstructured text (clinical notes) using large language models. NLP approaches that read physician notes could capture risk signals — like a family history mentioned in passing — that structured data fields miss.

The authors recommend that future studies consistently report exclusion windows, use explainable AI, validate on geographically diverse external datasets, and clearly distinguish between PDAC and other pancreatic cancer subtypes. Standardizing these practices would make the field's progress much easier to assess.

TL;DR: Near-term applications are possible, but the field needs better reporting standards, explainable AI, and large language models to unlock the full potential of EHR-based risk prediction.
Citation: Open Access, 2024. Available at: PMC11296923.