Construction of a prognostic model for endometrial cancer related to programmed cell death using WGCNA and machine learning algorithms

Front Immunol 2025 AI 5 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Programmed Cell Death and Cancer Prognosis

Programmed cell death (PCD) refers to regulated processes by which cells deliberately destroy themselves in response to biological signals. Unlike accidental cell death, PCD follows specific molecular pathways and includes well-known forms like apoptosis (the classical self-destruct mechanism) as well as more recently discovered types including ferroptosis (iron-dependent death), pyroptosis (inflammatory death), and over a dozen other mechanisms.

In cancer, PCD plays a contradictory role. It can suppress tumors by eliminating cells that have accumulated dangerous mutations. But tumor cells frequently hijack or disable PCD pathways, allowing them to survive, proliferate unchecked, and resist chemotherapy. Understanding which PCD genes are active or silenced in a patient's tumor can reveal why some patients respond to treatment while others do not.

For endometrial cancer specifically, the relationship between PCD-related genes and prognosis remains poorly characterized. This study used data from 543 UCEC (uterine corpus endometrial carcinoma) patients in The Cancer Genome Atlas (TCGA) to build the first comprehensive PCD-based prognostic model for this cancer type, incorporating 18 different forms of cell death and 1,548 PCD-associated genes.

TL;DR: Programmed cell death genes are dysregulated in endometrial cancer; this study built the first prognostic model using 1,548 PCD-related genes from 18 distinct cell death pathways.
Pages 2-5
Subtyping and Gene Prioritization Pipeline

The researchers began by identifying 217 PCD-related differentially expressed genes (PCD-DEGs) in TCGA endometrial cancer patients compared to normal tissue. These genes were then used to perform consensus clustering, which divided the 543 patients into two molecular subtypes: C1 (n=322) and C2 (n=221). The C2 subtype had significantly worse survival and a higher degree of immune suppression.

Weighted Gene Co-expression Network Analysis (WGCNA) was applied to identify which genes were most tightly linked to the C2 poor-prognosis subtype. This identified a 'blue module' of 931 genes with strong correlation to the C2 cluster. The intersection of this module with the original 217 PCD-DEGs yielded 31 candidate genes for further analysis.

To select the most informative genes from these 31 candidates, five machine learning algorithms were applied independently: Boruta (random forest-based), XGBoost (gradient boosting), Gaussian Mixture Models (GMM), Support Vector Machine with Recursive Feature Elimination (SVM-RFE), and Random Forest. Only genes selected by all five methods were retained, yielding three core genes: SRPX, NT5E, and ATP6V1C2.

TL;DR: Patients were clustered into two survival subtypes; WGCNA and five machine learning algorithms converged on three core PCD genes - SRPX, NT5E, and ATP6V1C2 - for model construction.
Pages 5-11
Risk Score Model Performance

The three core genes were used to build a LASSO-Cox risk score model with the formula: risk score = (-0.068 x ATP6V1C2) + (0.274 x SRPX) + (-0.066 x NT5E). Patients were divided at the median risk score into high-risk and low-risk groups. Higher risk scores were strongly associated with worse overall survival across training, testing, and combined TCGA cohorts.

ROC curve analysis confirmed strong predictive performance: the model accurately predicted 1-year, 3-year, and 5-year survival, although performance was slightly weaker at the 3-year mark. A nomogram combining the risk score with clinical variables (age, stage, BMI, grade) achieved AUC values of 0.708 at 1 year, 0.702 at 3 years, and 0.737 at 5 years for overall survival prediction.

Expression levels of the three genes were validated experimentally using RT-qPCR in tissue samples from 20 EC patients and 20 controls. SRPX was decreased in cancer tissue compared to controls, while ATP6V1C2 and NT5E were elevated - patterns consistent with the model's risk coefficients and confirming the clinical relevance of these genes in actual patient samples.

TL;DR: The LASSO-Cox risk model using SRPX, NT5E, and ATP6V1C2 successfully stratified patient survival, with RT-qPCR validating the predicted expression patterns in clinical tissue samples.
Pages 12-15
Immune Landscape and Drug Sensitivity

High-risk patients showed elevated TIDE scores - a measure of tumor immune evasion potential - and lower predicted immunotherapy response rates (22% vs. 52% in low-risk patients). Immune cell infiltration analysis found more immunosuppressive M2 macrophages and activated mast cells in high-risk tumors, likely contributing to the hostile immune environment that enables tumor escape.

High-risk patients also had lower tumor mutation burden (TMB) and fewer diverse mutations, which appears paradoxical - typically high mutation loads signal stronger immune recognition. However, the combination of low TMB and high TIDE score suggests these tumors avoid immune attack through microenvironmental suppression rather than simply having fewer targets for immunity to recognize.

Drug sensitivity analysis identified treatment differences between risk groups: high-risk patients showed greater predicted sensitivity to Dactolisib, Gemcitabine, Camptothecin, and Luminespib, suggesting aggressive chemotherapy may be more appropriate for this group. Low-risk patients were better predicted to respond to targeted agents with lower toxicity profiles such as Dasatinib and BI-2536.

TL;DR: High-risk patients have immunosuppressive microenvironments with low immunotherapy response rates and benefit from more aggressive chemotherapy, while low-risk patients may respond better to targeted agents.
Pages 17-19
Clinical Value and Limitations

The PCD-based prognostic model represents the first study to integrate 18 forms of programmed cell death into a UCEC prognostic framework - a substantially broader scope than previous models that examined only one or two PCD types. This comprehensiveness may explain the strong predictive performance, as cancer cells exploit multiple death-evasion mechanisms simultaneously.

The model's three core genes - SRPX (tumor suppressor), NT5E (CD73, an immune regulator), and ATP6V1C2 (V-ATPase subunit) - each have distinct roles in cancer biology. Their combination captures both cell-intrinsic (proliferation, stress resistance) and immune microenvironment dimensions of EC prognosis, explaining the model's broad predictive power.

Limitations include reliance on a single TCGA dataset for primary model building and a small RT-qPCR validation cohort of 20 patients. An observed discrepancy in NT5E expression direction between the TCGA dataset and RT-qPCR validation is explained by the stage composition difference - TCGA had more advanced-stage patients while the RT-qPCR cohort was predominantly early-stage, and NT5E expression appears to change with disease progression. Multi-center prospective validation is the critical next step.

TL;DR: The model's three core genes capture both cancer-intrinsic and immune microenvironment dimensions of EC prognosis, but requires multicenter prospective validation before clinical adoption.
Citation: Open Access, 2025. Available at: PMC12129963.