Every cell relies on the endoplasmic reticulum (ER) - a network of membranes inside the cell - to fold newly made proteins into their correct shapes. When this process goes wrong and proteins pile up in an unfolded or misfolded state, the cell enters a state of endoplasmic reticulum (ER) stress. The cell responds through a set of survival and stress-response pathways collectively called the unfolded protein response.
In cancer, ER stress is a double-edged sword. Mild ER stress can help tumor cells survive the harsh conditions of a rapidly growing tumor - including low oxygen, nutrient scarcity, and DNA damage - by activating protective responses. But when ER stress becomes severe, it can trigger programmed cell death. Cancer cells exploit ER stress pathways to their advantage, using them to resist chemotherapy, promote invasion, and suppress immune responses.
Despite accumulating evidence linking ER stress to several cancer types, its specific role in endometrial cancer (UCEC) had not been systematically characterized with modern multi-omics approaches. This study combined gene expression analysis from multiple public databases, network analysis, and multiple machine learning algorithms to identify which ER stress genes matter most for endometrial cancer prognosis, then validated findings in actual patient tissue samples.
The study assembled gene expression data from multiple sources: the TCGA-UCEC cohort (579 samples), and three GEO datasets (GSE17025, GSE63678, GSE115810) merged together for discovery, plus GSE106191 as an external validation set. ER stress-related genes (ERGs) were drawn from the GeneCards database - only those with high relevance scores were included.
Weighted Gene Co-expression Network Analysis (WGCNA) was then applied separately to the TCGA and merged GEO datasets. WGCNA is a powerful technique that groups genes into modules based on how consistently they rise and fall together across samples, then identifies which modules correlate most strongly with clinical features like survival status. The genes from the top modules in each dataset were intersected with each other and with the ER stress gene list, yielding 94 differentially expressed ER stress-related genes (DEERGs) common to both datasets.
To narrow from 94 to the most important prognostic genes, four machine learning algorithms were applied in parallel: LASSO regression, XGBoost, SVM-RFE (Support Vector Machine - Recursive Feature Elimination), and Random Forest. Each algorithm independently selected a different number of candidate genes based on different mathematical approaches. Only genes selected by all four algorithms were accepted as hub genes - a stringent requirement that reduces the chance of a gene being selected by chance rather than true biological relevance.
Applying consensus clustering to 544 UCEC patients based on the expression of 94 DEERGs produced two optimal subtypes. Cluster C1 (n=252) had significantly better survival than C2 (n=292), which showed shorter overall survival and progression-free survival. The PCA plot confirmed that the two clusters were transcriptomically distinct, occupying clearly separated regions in multidimensional expression space.
The tumor microenvironments were strikingly different. C1 had higher immune scores (more immune infiltration) but lower tumor purity, while C2 had higher tumor purity and lower immune activation. CIBERSORT immune cell analysis found that C2 had more CD4 memory activated T cells, M1 and M2 macrophages, and dendritic cells, but also more regulatory T cells (Tregs) that suppress anti-tumor immunity - a pattern consistent with a complex immunosuppressive environment where immune cells are present but dysfunctional.
Immune checkpoint gene expression was higher in C2: PD-L1, CTLA4, TIGIT, LAG3, and others were all more expressed in the worse-prognosis cluster. HLA gene expression was also generally higher in C2. This paradox - more immune activity but worse outcomes - points to simultaneous immune activation and suppression in C2, where immune cells are recruited but then suppressed, allowing cancer cells to escape destruction. This type of microenvironment may actually respond to immunotherapy, since the suppressive molecules are the targets of checkpoint inhibitors.
The intersection of all four machine learning algorithms identified exactly four hub genes: MYBL2, RADX, RUSC2, and CYP46A1. In UCEC tumors, MYBL2 was upregulated while the other three were downregulated compared to normal endometrial tissue. These bioinformatic findings were then directly confirmed in actual patient samples.
Real-time PCR on 20 UCEC and 20 normal tissue samples from a Chinese hospital showed the same pattern: MYBL2 mRNA was significantly higher in cancer, while RADX, RUSC2, and CYP46A1 mRNA were significantly lower. Immunofluorescence staining of proteins from the same tissue samples confirmed that the protein-level expression matched the mRNA findings - an important step because mRNA and protein levels do not always agree.
MYBL2 is a transcription factor involved in cell cycle regulation, and its elevated expression has been linked to poor prognosis in bladder cancer and liver cancer. RADX is a DNA stability protein; its loss can confer resistance to chemotherapy. CYP46A1 is involved in cholesterol metabolism and has been implicated in glioblastoma and colorectal cancer. RUSC2 is a structural domain protein linked to lung cancer progression. None of these had been previously characterized specifically in endometrial cancer.
A risk score was calculated for each patient: risk score = (0.113 x MYBL2) + (0.070 x RADX) + (0.030 x RUSC2) + (0.173 x CYP46A1). Patients above the median score were labeled high-risk. In the training set, high-risk patients had significantly worse overall survival. The same was true in the testing set and in the full TCGA cohort.
The AUC values for the risk score alone ranged from 0.562 to 0.688 at 1-, 3-, and 5-year survival in the training cohort, and 0.636 to 0.764 in the testing cohort. These are modest but meaningful prediction values. When the risk score was combined with clinical variables (age, stage, grade) into a nomogram - a visual prediction tool that doctors can use - performance improved substantially, with AUC values of 0.796, 0.804, and 0.844 at 1, 3, and 5 years respectively.
Decision curve analysis (DCA) confirmed that the nomogram provided positive clinical net benefit for 3- and 5-year survival prediction across a range of decision thresholds, meaning that using this tool to guide treatment decisions would, on average, produce better outcomes than either treating everyone or treating no one regardless of risk score.
Using the oncoPredict package against the GDSC drug sensitivity database, the researchers estimated how sensitive high-risk versus low-risk tumors would be to a panel of drugs. The analysis revealed clearly different drug sensitivity profiles between the two groups.
High-risk patients showed lower IC50 values (meaning they require less drug to kill cells) for ibrutinib (a BTK inhibitor originally used in blood cancers), cediranib (an anti-angiogenesis drug targeting VEGF receptors), fulvestrant (an estrogen receptor antagonist), and teniposide (a topoisomerase inhibitor). These drugs may therefore be more effective in high-risk endometrial cancer patients.
Low-risk patients were predicted to be more sensitive to dactinomycin (an older cytotoxic antibiotic), docetaxel (a chemotherapy taxane), selumetinib and trametinib (both MEK inhibitors used in targeted therapy). The divergence in predicted drug sensitivity suggests that risk stratification based on this model could one day help clinicians select which drug is most likely to be effective for a given patient, rather than applying a one-size-fits-all approach.