Soft tissue sarcomas (STS) are a rare but heterogeneous group of malignant tumors arising primarily from mesoderm-derived tissues, including muscle, fat, mucous tissue, mesothelium, blood vessels, and lymphatic vessels. There are more than 50 recognized subtypes distributed across 19 histological categories, and STS collectively account for approximately 1% of all adult malignancies but a disproportionately higher share of cancer-related deaths in younger patients. The rarity and diversity of STS subtypes create significant challenges for prognosis, treatment selection, and clinical trial design.
The role of programmed cell death: Cell death occurs through two broad mechanisms: accidental cell death (ACD), which is an uncontrolled biological process, and programmed cell death (PCD), which is genetically regulated. PCD encompasses at least 12 distinct modes, including apoptosis, necroptosis, ferroptosis, pyroptosis, autophagic cell death, parthanatos, entosis, netosis, lysosome-dependent cell death, oxeiptosis, alkaliptosis, and anoikis. Each of these modes operates through distinct molecular pathways and interacts differently with the tumor immune microenvironment. The balance and activity of these PCD pathways in tumor cells can shape both tumor progression and the immune landscape, making PCD-related gene expression a potentially powerful prognostic source.
The clinical gap: Despite the biological importance of PCD in oncology, no validated multi-PCD-pattern prognostic signature existed for STS prior to this study. Existing prognostic models for STS rely on clinical variables such as tumor grade, size, depth, and surgical margin status. These models provide only coarse risk stratification and cannot capture the molecular heterogeneity within clinical risk groups. Machine learning approaches that integrate transcriptomic data across multiple PCD pathways offer a route to more granular prognosis and, crucially, a biologically grounded basis for predicting immunotherapy response.
This 2025 study, published in Discover Oncology by Liu, He, and colleagues, set out to build a novel cell death-related signature (CDSig) for STS using a systematic framework of 96 machine learning algorithm combinations, validated across four independent cohorts.
The study began by compiling a curated set of 1,078 PCD-related genes from prior literature, covering all 12 recognized PCD modes. Transcriptomic data and paired clinical information were drawn from three major databases: The Cancer Genome Atlas (TCGA), the Therapeutically Applicable Research to Generate Effective Treatments (TARGET) program, and the Gene Expression Omnibus (GEO), specifically datasets GSE21250 and GSE71118. The TCGA sarcoma cohort served as the training set, while TARGET, GSE71118, and GSE21050 functioned as independent testing sets. Normal tissue expression data were obtained from the Genotype-Tissue Expression (GTEx) database to enable tumor-versus-normal differential analysis.
Differential expression and prognostic filtering: Differential expression analysis between STS and normal control tissues was performed using the edgeR R package, with filtering criteria of P < 0.05 and absolute log2 fold change greater than 1. This initial filter identified 439 PCD-related differentially expressed genes (DEGs). Principal component analysis (PCA) and Uniform Manifold Approximation and Projection (UMAP) were then applied to confirm the discriminative capacity of these DEGs between tumor and normal tissues. Univariate Cox regression further narrowed the list to 84 prognostically significant DEGs, which served as the input feature set for the machine learning pipeline.
The 96-algorithm combination framework: The core methodological innovation of this study was the construction and systematic comparison of 96 machine learning frameworks by combining 10 distinct algorithms: Random Survival Forest (RSF), Elastic Net (Enet), Lasso, Ridge, Stepwise Cox, CoxBoost, Cox Partial Least Squares Regression (plsRcox), Supervised Principal Component (SuperPC), Generalized Boosted Regression Modeling (GBM), and Survival Support Vector Machine (survival-SVM). Each framework applied an initial algorithm for feature selection to identify prognostic gene variables, followed by a second algorithm to compute an optimal prognostic score. Ten-fold cross-validation was used throughout model development to prevent overfitting.
Model selection criterion: Performance across the 96 frameworks was quantified using the concordance index (C-index) calculated simultaneously across all four cohorts (TCGA, TARGET, GSE21250, GSE71118). The model achieving the highest average C-index across all cohorts was selected as the final CDSig. The winning combination was RSF for initial feature selection followed by Supervised Principal Component (SuperPC) for signature construction. The resulting CDSig comprises 25 genes with the highest prognostic significance among the original 84 candidates.
The 25-gene CDSig was used to compute a continuous risk score for each patient, with patients stratified into high-CDSig and low-CDSig groups using the median as a cutoff. Kaplan-Meier survival analysis demonstrated that high-CDSig patients had significantly worse overall survival than low-CDSig patients in all four cohorts (TCGA, TARGET, GSE21250, GSE71118), with log-rank test p-values below 0.05 in each case. This consistent separation across heterogeneous cohorts collected from different institutions and time periods provides strong evidence for the signature's generalizability.
C-index comparison against clinical variables: The prognostic performance of CDSig was benchmarked against common clinical predictors including tumor grade, metastatic status, age, gender, and tumor size. In the TCGA cohort, CDSig exhibited the highest C-index among all evaluated features, outperforming every clinical variable tested. In the TARGET cohort, CDSig ranked second only to metastatic status, which is generally considered the single strongest prognostic factor in STS. This positioning demonstrates that CDSig adds prognostic information beyond what is already captured by clinical staging.
Time-dependent AUC and calibration: The time-dependent AUC of CDSig for 1-, 3-, and 5-year overall survival prediction was computed using the timeROC R package. AUC values demonstrated strong performance at each time point across all cohorts, with the calibration curves showing close agreement between predicted and observed survival probabilities. Decision curve analysis (DCA) further confirmed that using CDSig for clinical decision-making provides net benefit across a wide range of probability thresholds, supporting its practical utility beyond purely statistical performance metrics.
Cox regression validation: Univariate Cox regression confirmed that CDSig was significantly associated with overall survival (hazard ratio greater than 1 for high CDSig, p < 0.05) in both TCGA and TARGET cohorts. Crucially, multivariate Cox regression adjusting for age, gender, and metastatic status showed that CDSig remained an independent prognostic factor in both cohorts, indicating that its prognostic value is not simply a proxy for known clinical risk factors.
To translate the CDSig into a clinically usable tool, the authors constructed a nomogram that combines the CDSig score with standard clinical characteristics: gender, age, and metastatic status (primary metastasis or metastasis detected at follow-up). Nomograms are graphical tools that convert multivariable regression model results into individual probability estimates, making them practical for bedside use. The nomogram was built using the rms and regplot R packages based on the TCGA cohort and outputs the estimated probability of 1-, 3-, and 5-year overall survival for individual patients.
Nomogram calibration and discrimination: Calibration plots comparing predicted versus observed 1-, 3-, and 5-year survival probabilities showed strong agreement in the TCGA cohort, indicating that the nomogram's probability estimates are reliable rather than systematically over- or under-predicted. The AUC of the nomogram across all three time points exceeded that of CDSig alone and that of any individual clinical variable, confirming additive benefit from integrating molecular and clinical data. This is consistent with the well-established principle that combined molecular-clinical models outperform either source in isolation.
Decision curve analysis: Decision curve analysis (DCA) was performed to evaluate the clinical utility of the nomogram at 1-, 3-, and 5-year survival prediction. DCA compares the net clinical benefit of a model against the extreme strategies of treating all patients or treating no patients across a range of decision thresholds. The nomogram showed positive net benefit across a clinically relevant range of threshold probabilities, outperforming both extreme strategies and the individual components (CDSig alone, clinical variables alone). This confirms that the combined nomogram would improve clinical decision-making in a real-world deployment.
To understand the biological underpinnings of CDSig risk stratification, the authors performed gene set enrichment analysis (GSEA) and functional annotation using the clusterProfiler R package, separately examining pathways enriched in the high-CDSig and low-CDSig groups. The top 20 enriched biological processes in each group revealed strikingly different tumor biology, providing a mechanistic rationale for the observed survival differences.
High-CDSig group biology: Patients in the high-CDSig group, who have worse prognosis, showed enrichment for processes related to development, cell adhesion, and morphogenesis, including embryonic morphogenesis, growth factor signaling, and cell-cell differentiation pathways. This profile is consistent with a more stem-like, de-differentiated tumor biology associated with aggressive behavior, resistance to therapy, and greater capacity for invasion and metastasis.
Low-CDSig group biology: In contrast, the low-CDSig group, with better prognosis, showed enrichment for immune-related and inflammatory pathways. This suggests that tumors with lower CDSig scores are characterized by more active immune engagement, with greater immune cell infiltration and inflammatory signaling. The biological distinction between the two groups provides mechanistic support for why low-CDSig patients appear more likely to benefit from immunotherapy.
Intratumor heterogeneity, proliferation, and TGF-beta: Three additional tumor microenvironment metrics were compared between CDSig groups. Intratumor heterogeneity scores, which measure the degree of clonal diversity within a tumor, were significantly higher in the high-CDSig group, consistent with greater genomic instability and poorer outcomes. Proliferation scores were also elevated in the high-CDSig group, reflecting faster tumor growth kinetics. TGF-beta response scores, which indicate the degree of immunosuppressive signaling via the TGF-beta pathway, did not differ significantly between groups, suggesting CDSig stratifies immune activity through mechanisms other than TGF-beta suppression.
The tumor immune microenvironment (TME) is a key determinant of both prognosis and response to immune checkpoint inhibitors. To characterize the relationship between CDSig and the STS immune landscape, the authors applied three complementary analytical methods: the ESTIMATE algorithm to compute immune and stromal scores, the CIBERSORT algorithm to estimate relative abundances of 22 immune cell types, and single-sample gene set enrichment analysis (ssGSEA) to score specific immune cell populations and functional states. Spearman correlations between CDSig scores and individual immune metrics were computed across the TCGA cohort.
Low-CDSig tumors have more active anti-tumor immunity: Across multiple metrics, the low-CDSig group showed a more immunologically active microenvironment. ESTIMATE analysis revealed higher immune scores and higher ESTIMATE scores (a composite measure of immune and stromal infiltration) in the low-CDSig group. CIBERSORT analysis showed significantly elevated infiltration of M1 macrophages, CD8+ T cells, resting dendritic cells, and resting mast cells in the low-CDSig group. All of these cell populations are associated with active anti-tumor immune responses, with CD8+ T cells and M1 macrophages being the most directly cytotoxic.
Immune checkpoint expression is elevated in low-CDSig tumors: Expression of CTLA-4 and PD-1, the two most clinically targeted immune checkpoints, was higher in the low-CDSig group. At first glance this might seem paradoxical, since checkpoint expression is associated with immune exhaustion. However, elevated PD-1 and CTLA-4 in a tumor with active CD8+ T cell infiltration reflects an inflamed but suppressed immune microenvironment, which is precisely the context in which checkpoint inhibitors are most effective. This pattern predicts that low-CDSig patients have "hot" tumors that are actively engaging immunity but being held back by checkpoint suppression, making them the most likely candidates for response to anti-PD-1 or anti-CTLA-4 therapy.
High-CDSig tumors display immune exclusion features: By contrast, the high-CDSig group had higher tumor purity scores and higher proportions of M0 macrophages, resting NK cells, and mast cells, alongside lower leukocyte fractions and lower lymphocyte infiltration signature scores. This immunologically "cold" profile is associated with immune exclusion or immune desert tumor phenotypes, which respond poorly to checkpoint inhibitors and represent a major unmet therapeutic challenge in STS.
A central translational goal of this study was to test whether CDSig can predict which STS patients are likely to respond to immunotherapy and to specific chemotherapy agents. Two independent computational methods were used to assess immunotherapy response: the Tumor Immune Dysfunction and Exclusion (TIDE) algorithm, which models the probability of immune escape from T-cell killing, and the Subnetwork Mappings in Alignment of Pathways (Submap) method, which matches patients to transcriptomic profiles of known responders versus non-responders in clinical immunotherapy trial datasets.
TIDE analysis: Using the TIDE algorithm, low-CDSig patients in both the TCGA and TARGET cohorts showed a significantly greater proportion of predicted immunotherapy responders (TIDE score flagged as "true") compared to high-CDSig patients. The TIDE algorithm specifically models two mechanisms of immune failure, T-cell exclusion and T-cell dysfunction, and flags patients as likely responders when neither mechanism is strongly active. The low-CDSig group's elevated immune infiltration and checkpoint expression profile is consistent with TIDE's prediction of better immunotherapy responsiveness.
Submap analysis for anti-PD-1 and anti-CTLA-4: Submap analysis, which directly compares tumor transcriptomes against those of known responders and non-responders in published immunotherapy trial cohorts, confirmed that low-CDSig patients were significantly more likely to respond to both anti-PD-1 therapy and anti-CTLA-4 therapy (p < 0.05 for both). This convergent evidence from two independent prediction methods substantially strengthens the conclusion that CDSig stratifies immunotherapy sensitivity.
Drug sensitivity prediction: The GDSC (Genomics of Drug Sensitivity in Cancer) database was used to estimate IC50 values for common STS chemotherapy agents across CDSig groups. Patients in the low-CDSig group had significantly lower predicted IC50 values for doxorubicin, axitinib, cisplatin, and camptothecin compared to high-CDSig patients. A lower IC50 indicates that the drug is effective at a lower concentration, meaning low-CDSig patients may derive more benefit from these agents. This finding, combined with the immunotherapy predictions, suggests CDSig could guide treatment allocation between immunotherapy and chemotherapy in a precision oncology framework.
To provide experimental support for the computational CDSig, the authors selected the top 8 genes from the 25-gene signature and measured their mRNA expression in STS cell lines using RT-qPCR. The cell lines tested included SW872 (liposarcoma), SYO-1 (synovial sarcoma), and 005R, with human skin fibroblast (HSF) cells serving as a normal tissue control. The selected genes span multiple PCD modes and include PLCG1, BAG6, SQLE, ACSF2, GSS, MYBBP1A, PPP2R1B, and YWHAE.
RT-qPCR findings: The PCR results confirmed differential expression of all 8 genes across the STS cell lines relative to the normal HSF control. PLCG1, BAG6, and SQLE showed elevated expression in the sarcoma lines, while ACSF2, GSS, MYBBP1A, PPP2R1B, and YWHAE displayed reduced expression. This pattern aligns with the CDSig's underlying biology: genes linked to cell proliferation, survival, and immune evasion tend to be upregulated in high-risk STS cells, while genes associated with metabolic regulation and cell death execution are downregulated. PLCG1 in particular has been linked to poor prognosis in IDH wild-type low-grade glioma, and SQLE (squalene epoxidase) has established roles in cholesterol biosynthesis and tumor lipid metabolism.
Study limitations: The authors acknowledge several important constraints. All four datasets used for training and validation are derived from retrospective, single-center or database-compiled cohorts, meaning the CDSig has not been prospectively validated in a clinical trial setting where patient selection, treatment protocols, and data collection are controlled. The limited availability of clinical and molecular data in public repositories restricted inclusion of additional prognostic variables such as detailed chemotherapy regimens, surgical margins, and tumor depth, all of which are standard STS prognostic factors. Additionally, single-center retrospective designs are prone to spectrum bias and may not capture the full heterogeneity of STS presentations seen at large sarcoma referral centers.
Future directions: The authors call for prospective, multi-center validation cohorts to establish the real-world clinical utility of CDSig. Integration of CDSig with digital pathology and radiomics features extracted from MRI or CT imaging could further improve prognostic granularity. The 25 CDSig genes also represent potential therapeutic targets: small molecules or biologics that modulate PCD pathways, particularly ferroptosis inducers, necroptosis activators, and pyroptosis-stimulating agents, are an active area of investigation and could be prioritized for STS based on the CDSig framework. Ultimately, a prospective randomized study using CDSig to guide treatment allocation between immunotherapy and chemotherapy would be the definitive test of its clinical value.