Development of a Machine Learning-Based Predictive Model for Lung Metastasis in Patients With Ewing Sarcoma

Frontiers in Medicine 2022 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Why Predicting Lung Metastasis in Ewing Sarcoma Is Clinically Critical

Ewing sarcoma (ES) is an aggressive bone and soft tissue malignancy that predominantly affects children and adolescents, representing the second most common primary bone malignancy in this age group and accounting for roughly 5% of all pediatric cancers. The disease carries a high propensity for both local recurrence and distant metastasis, and despite advances in multimodal treatment combining chemotherapy, surgery, and radiation, the 5-year overall survival rate for patients with distant metastasis remains only 20-45%. Data from a landmark study of 975 ES patients illustrated this disparity starkly: localized disease carries a 5-year survival of 70% and a 5-year relapse-free survival of 55%, whereas patients with distant metastasis face 5-year survival of just 33% and relapse-free survival of 21%.

The lung as the primary metastatic site: Among the various distant sites to which ES can spread, the lung is by far the most common. Standard practice for detecting pulmonary nodules involves CT scanning of the chest, but this approach has significant drawbacks: high cost, cumulative radiation exposure, and a detection rate for metastatic nodules that reaches only approximately 20-25% of ES patients. This detection gap means that a substantial proportion of patients who will eventually develop overt lung metastasis are not identified early enough to prompt treatment intensification or alternative strategies.

The case for predictive modeling: Because CT-based surveillance alone is insufficient to identify all at-risk patients, predictive models capable of estimating lung metastasis (LM) probability at diagnosis from routinely available clinical and pathological variables would represent a meaningful advance. The authors of this 2022 study published in Frontiers in Medicine set out to fill this gap by developing and externally validating machine learning (ML)-based predictive models, then deploying the best-performing model as a freely accessible online tool for clinicians.

The study is notable for its multicenter design: training data came from the large SEER (Surveillance, Epidemiology, and End Results) database covering approximately 26% of the US population, while external validation used real patient data from four geographically distinct hospitals in China, providing a cross-population stress test of model generalizability.

TL;DR: Ewing sarcoma carries a 5-year survival of only 20-45% with distant metastasis, compared to 70% for localized disease. CT detects lung metastasis in only 20-25% of ES patients. This study trains 6 ML models on 929 SEER cases and validates them externally on 51 multicenter cases to predict lung metastasis at diagnosis, then deploys the best model as a web tool.
Pages 2-3
Study Design: SEER Training Cohort, Multicenter Validation, and Six ML Algorithms

The study enrolled 980 ES patients in total, dividing them into a training group (n = 929, sourced from the SEER database, 2010-2016) and an external validation group (n = 51, sourced from four Chinese medical institutions: Liuzhou People's Hospital, Second Affiliated Hospital of Jilin University, Xianyang Central Hospital, and Second Affiliated Hospital of Dalian Medical University, from 2010 to 2018). All selected cases required complete clinicopathological data and no additional primary tumors. The SEER cohort was filtered using ICD-O-3/WHO 2008 morphology code 9260d to identify bone-originating ES specifically. The retrospective nature of the design and use of de-identified registry data did not require ethics committee approval.

Variable selection: Fourteen demographic and clinicopathological variables were extracted for both cohorts: race, age, sex, primary site (axial bone, limb bone, or other), laterality, T-stage, N-stage, M-stage, surgical treatment, radiation, chemotherapy, bone metastasis status, and survival time. Variables in the SEER database were harmonized with the multicenter dataset by classifying all Chinese patients under the race category "other." Surgical treatment, radiation, and chemotherapy were each coded as yes or no, without granular details on specific approaches or agents.

Six ML algorithms applied: The authors trained six distinct ML classifiers independently on the SEER cohort: Random Forest (RF), Logistic Regression (LR), Extreme Gradient Boosting (XGBoost, or XGB), Gradient Boosting Machine (GBM), Multilayer Perceptron (MLP, a feedforward neural network), and Decision Tree (DT). All were implemented in Python 3.8. To mitigate overfitting during training, 10-fold cross-validation was applied, meaning each model was trained on 90% of the data and tested on the remaining 10%, cycling through all 10 folds, with the average AUC across folds used as the primary performance metric.

Feature importance analysis: To identify which variables drove prediction, the authors used permutation feature importance, a model-agnostic technique that evaluates each variable's contribution by measuring how much prediction performance degrades when that variable's values are randomly shuffled across 100 independent training simulations. Spearman correlation analysis with a correlation heat map was also applied to examine relationships between the four key retained variables, ensuring their independence as model inputs.

TL;DR: Training: 929 SEER patients (2010-2016). External validation: 51 patients from 4 Chinese hospitals (2010-2018). Six ML algorithms (RF, LR, XGB, GBM, MLP, DT) trained in Python 3.8 with 10-fold cross-validation. Performance assessed by AUC. Permutation feature importance (100 simulations) and Spearman correlation used to rank and verify independence of predictive variables.
Pages 3-4
Patient Demographics and the Profile of Lung Metastasis at Diagnosis

Across the full 980-patient cohort, lung metastasis was present at diagnosis in 185 cases (18.9%), a rate consistent with established epidemiological estimates for ES. The median age of the cohort was 22.25 years (SD = 16.3), reflecting the characteristic pediatric and young adult distribution of ES. Males comprised 57.5% (534/929) of the SEER training group, and over 85% were White, consistent with the known higher incidence of ES in individuals of European ancestry. Of the 929 SEER patients with data available for the LM vs. no-LM comparison, 175 (18.8%) had lung metastasis.

Statistically significant differences between groups: Comparing patients who developed LM (n = 175) versus those who did not (n = 754), statistically significant differences emerged for T-stage, N-stage, M-stage, surgical treatment, bone metastasis, and survival time (all p < 0.001). These variables formed the primary candidates for the multivariable models. In contrast, age, sex, race, primary site, laterality, radiation, and chemotherapy status did not differ significantly between the two groups, suggesting these factors alone are insufficient to stratify LM risk.

Training vs. validation group comparisons: The SEER training group and the multicenter validation group differed significantly in race distribution (by design, as all Chinese patients were coded as "other"), T-stage distribution, and radiation rates (22.8% vs. 43.1% with radiation, p = 0.001). Despite these differences, the overall LM rate was nearly identical: 18.8% in the SEER group and 19.6% in the multicenter group, suggesting comparable case mix in terms of the outcome of interest. These demographic and treatment differences between training and validation cohorts represent a realistic challenge to model generalizability and make the external validation results particularly informative.

Clinical patterns in the LM group: Bone metastasis co-occurred with lung metastasis at a notably elevated rate: 40.6% of patients with bone metastasis (56/138) also had LM, compared to only 15% of patients without bone metastasis (119/791). This synergistic metastatic pattern is clinically important and supports bone metastasis status as an independent predictive variable. Surgical treatment was significantly less common in the LM group (33.1% underwent surgery) than in the no-LM group (64.1%), which is likely explained by both the practical difficulty of resecting large or advanced tumors and the higher disease burden at presentation in patients who subsequently develop metastasis.

TL;DR: 185/980 patients (18.9%) had lung metastasis at diagnosis; median age 22.25 years. T-stage, N-stage, M-stage, surgery, bone metastasis, and survival time all differed significantly (p < 0.001) between LM and no-LM groups. Bone metastasis co-occurred with LM in 40.6% vs. 15% of patients without bone metastasis. LM rate was nearly identical between SEER (18.8%) and multicenter validation (19.6%) cohorts.
Pages 4-5
Logistic Regression Identifies Five Independent Predictors of Lung Metastasis

Before building the ML models, the authors conducted conventional univariate and multivariate logistic regression analysis to identify which variables independently predicted LM. Univariate analysis identified five variables with significant associations (p < 0.05): survival time, T-stage, N-stage, surgical treatment, and bone metastasis. All five were then entered into the multivariate logistic regression model.

Multivariate regression results: T-stage emerged as a strong independent predictor, with escalating odds ratios for higher stages compared to T1. Specifically, T2 carried an OR of 2.70 (95% CI: 1.690-4.317, p < 0.001), T3 an OR of 4.04 (95% CI: 1.773-9.194, p < 0.01), and TX (unknown/unclassified T-stage) an OR of 3.15 (95% CI: 1.778-5.566, p < 0.001). Lymph node involvement (N1 stage) showed the single strongest multivariate odds ratio at 5.10 (95% CI: 3.048-8.540, p < 0.001), more than quintupling the odds of LM relative to N0. Bone metastasis independently increased LM odds by a factor of 1.69 (95% CI: 1.090-2.605, p < 0.05).

Protective and neutral factors: Surgical treatment was a protective factor against LM, with an OR of 0.45 (95% CI: 0.309-0.658, p < 0.001), meaning patients who underwent surgery were significantly less likely to develop lung metastasis. Longer survival time was also associated with lower LM odds (OR = 0.988, 95% CI: 0.979-0.997, p < 0.01). The authors note, however, that survival time is not meaningfully known at the time of initial diagnosis and is therefore impractical as a real-time clinical predictor, which is precisely why it was excluded from the final ML models. NX (unknown lymph node status) did not reach significance in multivariate analysis (OR = 1.41, p = 0.302), suggesting that its apparent univariate association was confounded by other variables.

Rationale for the four-variable ML model: Based on these findings, and prioritizing variables that are clinically available at the time of initial ES diagnosis, the ML models were built using four variables: T-stage, N-stage, surgical treatment status, and bone metastasis. Survival time was excluded for practical reasons despite statistical significance. This four-variable structure strikes a balance between model parsimony and predictive completeness, using only information a treating clinician would have in hand at diagnosis.

TL;DR: Multivariate logistic regression identified 5 independent predictors. Strongest: N1 stage (OR = 5.10), T3 stage (OR = 4.04), TX stage (OR = 3.15), T2 stage (OR = 2.70), bone metastasis (OR = 1.69). Surgery was protective (OR = 0.45). Survival time excluded from ML models as impractical at diagnosis. Final ML models used T-stage, N-stage, surgery, and bone metastasis.
Pages 5-6
Random Forest Outperforms Five Competing Algorithms in Internal and External Validation

Six ML classifiers were trained on the 929-patient SEER cohort and internally evaluated using 10-fold cross-validation. The Random Forest (RF) model achieved the highest average AUC of 0.775 across the 10 folds, followed by the other algorithms in descending order. The Multilayer Perceptron (MLP) showed the lowest cross-validation performance among the six. This internal validation established RF as the leading candidate before any external data were examined.

External validation results: When all six models were applied to the 51-patient multicenter validation cohort from the four Chinese hospitals, the RF model again ranked first with an AUC of 0.705. The AUC range across the six models in external validation was 0.585 to 0.705, indicating meaningful variation in how well different algorithm types generalized to the unseen external data. The consistent top performance of RF in both internal and external settings provides more reliable evidence of model utility than an internal-only result would offer.

Why Random Forest performed best: The RF algorithm is an ensemble method that builds many decision trees on random subsets of the training data and features, then aggregates their predictions by majority vote. This "bagging" approach reduces the variance of individual trees and makes RF particularly robust to overfitting, which is a critical concern with small sample sizes and high-dimensional clinical data. RF also handles nonlinear relationships between variables and the outcome naturally, without requiring explicit variable transformation or interaction terms as logistic regression does.

Context for the AUC values: The external validation AUC of 0.705 is modest by the standards of some imaging-based or genomic AI models, but should be interpreted in context. The model uses only four simple clinical variables available at initial diagnosis, requires no imaging analysis or molecular testing, and is designed for a rare disease with a training dataset of under 1,000 cases. In this context, an AUC consistently above 0.70 in external validation represents a meaningful discriminative capacity that could supplement clinical judgment. The drop from internal AUC (0.775) to external AUC (0.705) is characteristic of models validated on demographically distinct populations and is within the range typically seen when SEER-trained models are applied to non-US cohorts.

TL;DR: RF model achieved internal 10-fold cross-validation AUC = 0.775 (best among 6 algorithms). In external validation on 51 multicenter cases, RF again led with AUC = 0.705. Overall external AUC range across all 6 models: 0.585-0.705. MLP performed worst internally. RF's ensemble bagging approach gave it superior generalizability.
Pages 6-7
Surgery, T-Stage, and N-Stage Consistently Drive Predictions Across All Models

To understand which of the four clinical variables contributed most to lung metastasis prediction, the authors applied permutation feature importance analysis across 100 independent training simulations for each model. In this technique, each variable's importance is quantified by measuring the drop in model AUC when that variable's values are randomly shuffled, thereby breaking any predictive relationship it has with the outcome. A variable whose shuffling causes a large performance drop is considered highly important; one whose shuffling causes little change contributes minimally.

Consistent variable rankings: Across all six ML models, surgery, T-stage, and N-stage ranked consistently in the top three positions for feature importance. Within the RF model specifically, the importance order was: surgery > T-stage > N-stage > bone metastasis. Bone metastasis consistently ranked fourth across the models. The consistency of this ordering across diverse algorithm types (tree-based, neural network-based, and linear) strengthens confidence that these rankings reflect genuine biological associations rather than artifacts of any single modeling approach.

Why surgery leads: The primacy of surgical treatment as the most important predictor is clinically interpretable from multiple angles. Surgery is not merely a treatment decision but also a marker of disease stage and tumor resectability: patients with larger, more advanced, or metastatically disseminated tumors are less likely to receive curative-intent surgery, meaning that absence of surgery often reflects underlying high-risk disease characteristics beyond what T-stage alone captures. Additionally, surgical pathology specimens enable more precise TNM staging and identification of surgical margin status, which carries prognostic weight. The authors highlight this dual diagnostic and therapeutic role of surgery as a key insight from the ML analysis.

Independence of variables: The Spearman correlation heat map confirmed no significant positive correlation between any pair of the four key variables. The only notable relationship was a negative correlation between surgery and the other three variables, which is biologically coherent: advanced T-stage, positive nodes, and bone metastasis all make curative surgery more difficult or inadvisable, so their co-occurrence with "no surgery" is expected. This lack of multicollinearity confirms that all four variables contribute independent, non-redundant information to the predictive models.

TL;DR: Permutation importance (100 simulations) ranked variables consistently across all 6 models: surgery first, then T-stage, N-stage, and bone metastasis fourth. RF-specific order: surgery > T-stage > N-stage > bone metastasis. Spearman correlation confirmed all four variables are independent inputs with no significant positive intercorrelation.
Pages 7-8
An Open-Access Web Calculator Translates the RF Model Into a Point-of-Care Tool

A core deliverable of this study is the translation of the best-performing RF model into a publicly accessible, web-based predictive tool hosted via the Streamlit platform. Clinicians can input a patient's four key variables, T-stage, N-stage, surgical treatment status, and bone metastasis status, and receive an estimated probability of lung metastasis. The tool was designed specifically to require no specialized computing knowledge from users and no installation of software beyond a standard web browser.

Clinical decision support context: The intended clinical use is not to replace imaging-based staging but to supplement it by providing a data-driven risk stratification at the time of initial ES diagnosis. High-probability predictions could prompt more aggressive imaging surveillance, earlier referral for pulmonary evaluation, or consideration of enrollment in clinical trials of treatment intensification. Low-probability predictions might reassure clinicians and patients in cases where imaging findings are ambiguous, potentially avoiding unnecessary procedures.

Practicality of the four-variable input set: A key strength of this tool design is that all four required inputs, T-stage, N-stage, bone metastasis status, and whether surgery was performed, are either already known at diagnosis from standard workup or are determined as part of the initial treatment decision. This means clinicians face no additional data collection burden to use the tool, a common barrier for more complex ML-based calculators that require specialized biomarkers, imaging processing, or molecular assays unavailable outside academic centers.

The accessibility of the web tool is particularly relevant for ES given its rarity. Community oncologists and institutions that manage only a handful of ES cases per year may have less intuitive calibration of lung metastasis risk than high-volume sarcoma centers. A validated, quantitative risk tool that distills information from nearly 1,000 training cases could meaningfully improve the consistency of risk communication and treatment planning across practice settings.

TL;DR: The best RF model was deployed as an open-access web calculator (Streamlit platform) using 4 clinically routine inputs: T-stage, N-stage, bone metastasis, and surgical treatment status. No specialized data collection is required. Target users include community oncologists managing rare ES cases who lack institutional experience to intuitively gauge LM risk.
Pages 8-9
Dataset Constraints, Missing Variables, and the Path Toward Better Predictive Tools

SEER database limitations: The SEER database, while large by rare-disease standards, does not capture several clinically important variables that are known to influence ES biology and metastatic potential. Specifically, the authors identify the following missing data: precise surgical technique and surgical margin status, tumor marker levels, evidence of vascular invasion, radiation dosage and field details, and chemotherapy regimen specifics (agents, cycles, and response). The absence of these variables means that predictive signals from chemosensitivity or margin-positive resection, both of which influence subsequent metastasis risk, could not be incorporated into the models.

Retrospective design and selection bias: The SEER-based training cohort is retrospective, which introduces selection bias: the patients included are those with sufficient documentation in the registry, potentially excluding patients who died early, transferred care, or had incomplete records. Retrospective datasets from large registries also reflect the care patterns of their collection era (2010-2016), and ES treatment protocols have continued to evolve, which may limit the applicability of a 2010-era model to current clinical populations.

Small external validation cohort: The multicenter validation cohort of 51 patients is small relative to the training set. This limits the statistical power to detect meaningful performance differences between models at external validation and makes confidence intervals around the external AUC estimates wide. The authors acknowledge that expanded multicenter studies with larger validation cohorts, particularly from diverse geographic regions, are needed to more rigorously characterize model performance and bias.

Future directions: The authors envision several improvements. Incorporating additional biological variables, such as tumor gene fusion status (EWSR1-FLI1, the canonical ES translocation), Ki-67 proliferation index, and serum lactate dehydrogenase, could substantially increase predictive power. Radiomics features extracted from CT or MRI of the primary tumor represent another untapped source of input variables. Larger, prospectively collected multicenter datasets, ideally leveraging international ES consortia data, would both improve model training and enable more definitive external validation. Finally, dynamic updating of risk estimates as treatment proceeds, incorporating interim imaging response and ctDNA data, would move the tool from a one-time baseline calculator to a longitudinal risk monitoring instrument.

TL;DR: Key limitations: SEER lacks surgical margin status, vascular invasion, tumor markers, and chemotherapy details. Retrospective design introduces selection bias. External validation cohort is only 51 patients. Future work should add molecular variables (EWSR1-FLI1, Ki-67, LDH), radiomics features, larger multicenter prospective datasets, and dynamic risk updating using treatment response data.