Pancreatic cancer has a 5-year survival rate of less than 8%, and up to 80% of patients already have distant metastases at the time of diagnosis. The liver is the most common site of metastasis, occurring in 37-41% of newly diagnosed cases.
Worse, even after apparently successful surgery, more than 60% of patients relapse with liver metastasis within two years. Accurate prediction of who is likely to develop liver metastasis would allow oncologists to tailor treatment plans — perhaps treating high-risk patients more aggressively from the outset.
This study used five different machine learning algorithms to build liver metastasis prediction models using a large, national population-based cancer database, comparing their performance to find the most reliable approach.
Data from 47,919 pancreatic cancer patients were retrieved from the Surveillance, Epidemiology, and End Results (SEER) database for the years 2010 to 2018. Of these, 15,909 (33.2%) had developed liver metastasis — providing a large, well-characterized dataset for model training.
Five machine learning algorithms were compared: logistic regression, extreme gradient boosting (XGBoost), support vector machine (SVM), random forest (RF), and deep neural network (DNN). This head-to-head comparison allowed identification of which approach is best suited for this prediction task.
After iterative feature selection, nine clinical variables were identified as the most predictive and used as inputs to the models. Model performance was assessed using receiver operating characteristic (ROC) curves, specificity, sensitivity, and risk stratification analysis.
The random forest (RF) model achieved the best performance, with an AUC of 0.871 on the training set and 0.832 on the test set. This demonstrates strong discriminative ability and good generalization from training to new data.
Risk stratification analysis divided patients into high-risk, intermediate-risk, and low-risk groups based on the model's predictions. The high-risk group had significantly higher liver metastasis rates and worse 5-year cancer-specific survival than the intermediate- and low-risk groups — validating that the model captures clinically meaningful distinctions.
Surgery, radiotherapy, and chemotherapy were found to significantly improve cancer-specific survival in both the high- and intermediate-risk groups, suggesting that the model's risk categories can help identify which patients benefit most from aggressive multimodal treatment.
Patients classified as high-risk for liver metastasis by the random forest model could be prioritized for systemic chemotherapy or neoadjuvant treatment before surgery — strategies that may reduce the risk of liver spread. Low-risk patients might be spared some of these aggressive interventions and their associated side effects.
The finding that treatment benefits differ by risk group suggests that the model could be used not just to predict outcomes but to match patients to the treatments most likely to benefit them — a step toward personalized oncology.
With only nine clinical variables needed — information routinely collected at diagnosis — this model could be easily implemented in clinical settings to support oncologists in making evidence-based decisions for every newly diagnosed patient.
This large population-based study demonstrates that machine learning — particularly random forest — can accurately predict liver metastasis risk in pancreatic cancer patients using routine clinical data, outperforming simpler statistical approaches.
The risk stratification framework provides a practical tool for oncologists to identify which patients need the most aggressive monitoring and treatment from the moment of diagnosis, rather than waiting for metastasis to appear.
Future prospective validation and refinement of these models — including the incorporation of molecular biomarkers — could further improve prediction accuracy and clinical utility in pancreatic cancer management.