Integrated machine learning based on cuproptosis and RNA methylation regulators to explore the molecular model of prostate cancer

J Cancer 2025 Machine Learning 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Two Newly Linked Biological Processes in Prostate Cancer

Prostate cancer (PCa) is the second most common cancer in men worldwide. While localized disease is often curable, advanced prostate cancer remains difficult to treat, with up to 30% of cases eventually progressing to castration-resistant forms that do not respond to standard androgen-blocking therapy.

This study investigates two emerging biological processes that appear to interact in prostate cancer. Cuproptosis is a newly described form of cell death triggered by copper ion accumulation inside cells. When copper builds up excessively, it disrupts cellular energy metabolism and causes proteins to clump together abnormally, killing the cell. Research has shown that cuproptosis-related genes influence tumor growth, metastasis, and immune suppression.

RNA methylation refers to chemical modifications added to RNA molecules, most commonly in the form of N6-methyladenosine (m6A). These modifications control how genes are expressed, affecting protein production, and have been linked to tumor development and spread in multiple cancer types. The m6A regulatory proteins are called RNA methylation regulators.

The crosstalk between cuproptosis and RNA methylation regulation had not previously been studied in prostate cancer. This study aimed to identify genes where these two biological systems intersect, collectively called cuproptosis-associated RNA methylation regulators (CARMRs), and use them to build a prognostic model.

TL;DR: This study explores how two newly described biological processes, copper-triggered cell death (cuproptosis) and RNA methylation, interact in prostate cancer to influence prognosis.
Pages 2-3
Identifying Key Genes Using Network Analysis

The researchers started with publicly available gene expression data from four independent prostate cancer cohorts: TCGA-PRAD (490 patients), GSE70768 (111 patients), GSE70769 (88 patients), and DKFZ (105 patients). All patients had information on recurrence-free survival (RFS), the primary outcome of interest.

A Pearson correlation analysis between known cuproptosis genes and RNA methylation regulators identified 25 CARMRs. To narrow these down to the most relevant ones for prostate cancer biology, the researchers applied single-sample gene set enrichment analysis (ssGSEA) to quantify how active the CARMR gene set was in each individual patient sample, generating a CARMR score.

The CARMR score was then fed into Weighted Gene Co-Expression Network Analysis (WGCNA), a technique that groups genes into modules based on how similarly they behave across samples. The module most strongly correlated with the CARMR score was identified, and its genes were overlapped with genes that were differentially expressed between prostate cancer and normal tissue. This process yielded 69 candidate CARMR genes for further analysis.

Univariate Cox regression analysis was used to screen these 69 genes for those significantly associated with recurrence-free survival, yielding 14 prognostic genes. After accounting for genes not measurable on available microarray platforms, 9 final genes were selected for model construction.

TL;DR: The study used gene correlation analysis, co-expression networks, and survival statistics across four independent patient cohorts to identify 9 key genes connecting cuproptosis and RNA methylation in prostate cancer.
Pages 3-5
Building the Optimal Prognostic Model from 101 Machine Learning Combinations

A major methodological strength of this study was its systematic approach to model selection. Rather than choosing a single machine learning algorithm arbitrarily, the researchers tested 10 different machine learning algorithms and all 101 pairwise combinations of those algorithms, using a 10-fold cross-validation framework to evaluate each combination's predictive performance.

The algorithms tested included random survival forest, elastic network, LASSO, Ridge regression, stepwise Cox, CoxBoost, partial least squares regression for Cox, supervised principal components, generalized boosted regression modeling, and survival support vector machine. Each was evaluated on the TCGA-PRAD training dataset, and performance was measured using the C-index, a measure of how well the model ranks patients by recurrence risk.

The best-performing combination was Ridge regression alone, which achieved the highest average C-index of 0.687. The final model assigned each patient a risk score based on the weighted sum of the 9 CARMR gene expression values, with coefficients from the Ridge regression determining how strongly each gene contributed.

The nine genes in the final model are: ATP5ME, BEND3, C4orf48, MACIR, SLC26A1, ENTPD5, ITGA2, LPP, and PIK3R1. Of these, SLC26A1 and C4orf48 had the two largest regression coefficients, making them the most influential predictors in the model.

TL;DR: By testing 101 combinations of 10 machine learning algorithms, Ridge regression was identified as optimal for building a 9-gene prostate cancer risk model with strong cross-validated performance.
Pages 5-7
High-Risk Patients Have Worse Prognosis Across All Four Datasets

Patients in each of the four cohorts were divided into high-risk and low-risk groups based on the median risk score. In all four datasets, patients in the high-risk group experienced significantly worse recurrence-free survival (p < 0.05 in all cohorts), validating the model's predictive power across different data sources and patient populations.

Time-dependent ROC analysis confirmed good model discrimination, with AUC values of 0.685, 0.695, and 0.632 for 1-year, 3-year, and 5-year recurrence prediction in TCGA-PRAD, and exceeding 0.70 in most validation cohorts. Both univariate and multivariate Cox regression analyses confirmed that the risk score was an independent prognostic factor, meaning it provided predictive value even after accounting for established clinical variables like Gleason score and tumor stage.

A nomogram incorporating the risk score, T stage, and Gleason score was constructed to give clinicians a practical tool for predicting 1-, 3-, and 5-year recurrence probability in individual patients. Calibration curves confirmed close agreement between predicted and observed outcomes, and decision curve analysis showed net clinical benefit compared to models using clinical variables alone.

Principal component analysis plots based on the 9 CARMR genes visually confirmed that high- and low-risk patients separated into distinct clusters in all four datasets, providing an intuitive visualization of the model's ability to stratify patients by biological risk.

TL;DR: The 9-gene CARMR model independently predicted prostate cancer recurrence in all four datasets, and combining it with Gleason score and tumor stage in a nomogram improved clinical prediction accuracy.
Pages 8-9
Biological Pathways and Drug Sensitivity Differ Between Risk Groups

Gene ontology and KEGG pathway analysis revealed that the CARMR genes are involved in neurological and muscle development pathways, including axon guidance, calcium signaling, cAMP signaling, and vascular smooth muscle contraction. This suggests these genes influence cell signaling and structural processes that may contribute to cancer progression.

Comparing pathway activity between high-risk and low-risk patients revealed that high-risk patients showed increased lipid metabolism and enzyme synthesis while low-risk patients had higher activity in amino acid metabolism and prostaglandin synthesis. This pattern is consistent with known metabolic reprogramming in aggressive prostate cancer, where lipid synthesis becomes a key energy source.

The high-risk group showed downregulation of androgen response, epithelial-mesenchymal transition, and inflammatory pathways. The reduced androgen response in high-risk patients is clinically significant because it suggests these tumors are less dependent on androgens, which aligns with the higher rate of progression to castration-resistant prostate cancer in high-risk patients.

Drug sensitivity analysis using the oncopredict algorithm found that high-risk and low-risk patients responded differently to common chemotherapy and targeted drugs. High-risk patients showed higher IC50 values for Doramapimod, Entospletinib, and JAK-8517 (less sensitive), while low-risk patients were less sensitive to Docetaxel, Epirubicin, Oxaliplatin, and 5-Fluorouracil. This differential drug response profile could guide personalized treatment selection.

TL;DR: High-risk CARMR patients show upregulated lipid metabolism, downregulated androgen signaling, and different drug sensitivity profiles compared to low-risk patients, pointing toward distinct treatment strategies for each group.
Pages 10-11
The Tumor Immune Microenvironment Differs by Risk Group

The tumor microenvironment (TME) consists of immune cells, blood vessels, and other non-cancer cells that surround the tumor and can either support or suppress it. Using multiple immune cell quantification algorithms (CIBERSORT, EPIC, TIMER, xCell), the researchers found that the high-risk group had higher abundance of memory T cells and helper T cells but lower abundance of cytotoxic immune cells that directly kill cancer cells.

This counterintuitive pattern suggests an immunosuppressive microenvironment in high-risk patients. Despite more immune cells being present, they appear to be in inactive or suppressive states rather than actively attacking the tumor. The abundance of most immune cell types decreased as the risk score increased.

Immune checkpoint analysis showed that most immune checkpoint genes were more highly expressed in low-risk patients. Immune checkpoints are regulatory proteins on immune cells that can be targeted by immunotherapy drugs like PD-1 inhibitors. Lower checkpoint expression in high-risk patients may indicate these patients are less likely to benefit from standard immunotherapy approaches.

Tumor mutational burden (TMB), which measures the total number of gene mutations in a tumor, was higher in the high-risk group and was independently associated with worse recurrence-free survival. When patients were subdivided by both risk score and TMB, the high-risk group consistently had worse prognosis regardless of TMB level, confirming the independent predictive value of the CARMR signature.

TL;DR: High-risk prostate cancer patients have an immunosuppressive tumor microenvironment with higher tumor mutational burden, lower immune checkpoint expression, and weaker anti-tumor immune responses.
Pages 11-12
Laboratory Validation of the Two Most Important Genes

The two genes with the largest regression coefficients in the model, C4orf48 and SLC26A1, were experimentally validated using quantitative real-time PCR (qRT-PCR) in human prostate cancer and normal prostate cell lines.

Both C4orf48 and SLC26A1 showed significantly higher expression in prostate cancer cell lines (LNCaP, C42, PC3) compared to the normal prostate cell line RWPE-1, confirming that these genes are biologically relevant to prostate cancer rather than statistical artifacts.

C4orf48 is a gene on chromosome 4 with previously limited characterization in prostate cancer. SLC26A1 is a sulfate transporter protein whose role in cancer metabolism and ion transport may link it to the copper-dependent metabolic disruptions central to cuproptosis pathways.

This laboratory validation provides an important bridge between the computational findings and actual cancer biology, supporting the relevance of these genes as real molecular drivers of prostate cancer risk rather than incidental statistical correlates.

TL;DR: Laboratory experiments confirmed that both C4orf48 and SLC26A1, the two genes with the largest influence in the prognostic model, are significantly overexpressed in prostate cancer cells compared to normal prostate cells.
Pages 12, 13, 15
A New Prognostic Tool for Personalized Prostate Cancer Management

This study successfully developed and validated a 9-gene prognostic signature based on the intersection of cuproptosis biology and RNA methylation regulation. The model was trained in one large cohort and independently validated in three additional datasets from different institutions, demonstrating its generalizability.

The clinical utility of this model extends beyond prognosis. By linking risk stratification to immune microenvironment characterization and drug sensitivity profiles, it offers a framework for personalized treatment decisions. High-risk patients identified by the model may be prioritized for more aggressive treatment, alternative therapeutic strategies, or enrollment in clinical trials targeting the specific pathways their tumors rely on.

The biological insights uncovered, particularly the connections between copper metabolism, RNA modification, lipid metabolic reprogramming, and immune suppression in aggressive prostate cancer, open new avenues for therapeutic development. Targeting CARMRs could disrupt multiple cancer-promoting pathways simultaneously.

Limitations include the retrospective nature of the datasets and the lack of direct clinical implementation data. Future prospective studies should validate whether risk stratification using this model can meaningfully guide treatment decisions and improve patient outcomes in real-world clinical settings.

TL;DR: The 9-gene CARMR signature provides clinically validated prostate cancer risk stratification and points toward new therapeutic targets at the intersection of copper metabolism and RNA modification biology.
Citation: Open Access, . Available at: PMC12171011.