Machine Learning-Driven Integration of Cancer Cell Phenotypes Predicts Cisplatin Sensitivity

Cancer Med 2025 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Extending Precision Medicine to Cisplatin

Cisplatin is one of the most widely used anticancer agents, administered to approximately 50 percent of all cancer patients, yet there are no validated biomarkers to predict which tumors will respond to it. In contrast to molecularly targeted agents where genetic alterations such as EGFR mutations predict treatment benefit, cisplatin's broad cytotoxic mechanism makes it difficult to identify predictive biomarkers using conventional genomic profiling approaches.

Without predictive biomarkers, cisplatin is prescribed based on cancer type and empirical evidence rather than individualized response prediction. This means a substantial fraction of patients receive cisplatin without benefit while being exposed to its well-documented toxicities including nephrotoxicity, ototoxicity, and peripheral neuropathy.

Cancer cells develop specific gene expression patterns that reflect their phenotypic state, including their propensity for DNA damage repair, apoptosis, and drug transport. These phenotypic transcriptional signatures, captured through bulk RNA sequencing, represent a dimension of cisplatin sensitivity prediction that goes beyond individual driver mutations to reflect the integrated biology of the tumor cell.

This study introduces CSP26G, a logistic regression model trained on 26 gene expression biomarkers identified through integrated analysis of differentially expressed genes and SHAP-value-based machine learning from cancer cell line data, validated experimentally in a cisplatin-resistant cell line and clinically in TCGA patient data.

TL;DR: Cisplatin is administered to half of all cancer patients without predictive biomarkers, and this study develops CSP26G, a 26-gene machine learning model that predicts cisplatin sensitivity from tumor RNA expression to enable more individualized treatment decisions.
Pages 2-4
Identifying Cisplatin-Sensitive and Resistant Cell Lines

IC50 values from two public databases, PRISM (pooled barcode cell line profiling) and GDSC2 (traditional cell viability assays), were integrated across 190 cancer cell lines using COSMIC IDs as common identifiers. Using both databases ensured that classified cell lines had consistent sensitivity profiles regardless of the methodology used to measure drug response, reducing noise from platform-specific artifacts.

Hierarchical clustering with Ward's method and Euclidean distance was applied to IC50 values from both databases simultaneously. Silhouette score analysis confirmed that four clusters provided the best separation. Clusters 2 (n=60, high IC50 in both databases) and 4 (n=55, low IC50 in both databases) were selected as the cisplatin-resistant and cisplatin-sensitive training groups, respectively, because they showed consistent profiles across measurement platforms.

The IC50 boundary between the sensitive and resistant groups was approximately 30 micromolar, which corresponds to the maximum plasma concentration of cisplatin achievable in humans. This alignment with physiologically relevant drug concentrations provides biological justification for the classification boundary and confirms clinical relevance of the machine-driven grouping.

Differentially expressed gene analysis was performed using PyDESeq2 on bulk RNA-seq data from the Cancer Dependency Map (DepMap) comparing clusters 2 and 4, using an adjusted p-value threshold of less than 0.05 and absolute log2 fold change greater than 1. The resulting DEG list formed one of two inputs to the final gene selection pipeline.

TL;DR: 190 cancer cell lines were classified into cisplatin-sensitive and resistant groups through hierarchical clustering of IC50 values from two databases, with the resistance boundary aligned to physiologically achievable cisplatin plasma concentrations.
Pages 4-5
SHAP-Based Feature Gene Identification

A LightGBM gradient-boosting classifier was trained to distinguish cisplatin-sensitive from cisplatin-resistant cell lines using all genes as input features, and SHAP (Shapley Additive Explanations) values were computed to quantify each gene's contribution to the classification. This approach captures gene-gene interaction effects that conventional DEG analysis cannot identify, because SHAP values reflect a gene's marginal contribution within the full high-dimensional gene expression context.

To ensure stable SHAP estimates, the model was trained and evaluated using 12 random data splits combined with 5-fold stratified cross-validation each time, yielding 60 independent SHAP computations per gene. Average SHAP values across these iterations identified 447 genes that consistently contributed to sensitivity prediction.

A Venn diagram analysis identified the overlap between DEG-selected genes and SHAP-selected genes, producing 55 candidate biomarker genes: 29 upregulated and 26 downregulated in the resistant versus sensitive comparison. This intersection approach selected genes that are both statistically differentially expressed and functionally influential in the machine learning model, reducing false positives from each analysis performed alone.

Recursive Feature Elimination with Cross-Validation (RFECV) was then applied to the 55-gene candidate set using logistic regression, iteratively removing the gene with the smallest absolute regression coefficient and evaluating specificity at each step. Peak specificity was achieved with exactly 26 genes, which became the final CSP26G biomarker set. Pairwise correlation analysis confirmed that no two genes in the final 26-gene set were strongly correlated, indicating each gene contributes independent predictive information.

TL;DR: Combining DEG analysis with SHAP-value machine learning identified 55 candidate genes, and recursive feature elimination optimized the final 26-gene biomarker set that maximizes specificity for cisplatin resistance prediction while maintaining independent contributions from each gene.
Pages 5-6
CSP26G Performance and External Validation

CSP26G achieved an AUC of 0.99 on the ROC curve, 0.99 on the precision-recall curve, and a sensitivity and specificity of 0.93 each on the held-out test dataset, demonstrating near-perfect discrimination between cisplatin-sensitive and resistant cell lines. Specificity was chosen as the primary optimization metric to minimize false positives (predicting resistance when a patient is actually sensitive), ensuring that truly responsive patients are not incorrectly denied cisplatin.

A cisplatin-resistant A549 lung cancer cell line (A549CR) was established by continuously exposing parental A549 cells to escalating cisplatin concentrations for 29 weeks, achieving an IC50 of 57.53 micromolar compared to 10.02 micromolar for parental A549 cells. RT-qPCR measurement of 26-gene expression in A549CR versus A549 cells showed a significantly positive correlation with the RNA-seq resistant/sensitive log2 fold changes from the training data, confirming that the gene expression pattern encoded in CSP26G is biologically recapitulated in experimentally acquired resistance.

Clinical validation in TCGA data from 407 stage IIA to IIIA NSCLC patients who received platinum-based adjuvant chemotherapy showed that CSP26G-predicted sensitive patients had significantly longer median overall survival (5.92 years) compared to CSP26G-predicted resistant patients (2.74 years). The survival curves diverged during the mid-term follow-up period, with ROC analysis reaching peak discrimination at 8 years (AUC = 0.76).

Multivariate Cox regression confirmed CSP26G score as the strongest independent predictor of survival after adjusting for age and TNM staging, outperforming conventional clinical variables. Importantly, CSP26G was not predictive in esophageal and ovarian cancer cohorts where cisplatin use was inconsistent, confirming that the model's predictive power is specific to populations receiving cisplatin treatment rather than reflecting a general cancer prognosis signal.

TL;DR: CSP26G achieved AUC 0.99 on test cell line data, was validated in experimentally derived cisplatin-resistant A549CR cells, and stratified TCGA NSCLC patients into groups with significantly different survival (5.92 vs 2.74 years median OS) when receiving cisplatin-based therapy.
Pages 8-9
Key Biomarker Genes and Mechanistic Insights

The 26-gene CSP26G set includes genes across diverse biological pathways, reflecting the multifactorial nature of cisplatin resistance that cannot be reduced to a single mechanism. Identified genes include the tumor suppressor EPHA7, the apoptosis regulator EGR3, the organic anion transporter SLCO5A1 that may influence drug cellular uptake, and STK32A, a putative mitotic checkpoint kinase.

The gene set also unexpectedly contains multiple immune response regulators including RORC and S1PR4, despite CSP26G being trained on cultured cell lines devoid of immune cells. This suggests these immune-associated genes may have intrinsic functions in cancer cell survival or drug response pathways beyond their canonical immunological roles, representing a biologically novel finding warranting further mechanistic investigation.

SLFN11, a well-established predictor of DNA-damaging agent sensitivity that induces TP53-independent apoptosis, was not retained in the final 26-gene set but was identified as a significant SHAP contributor among the 447-gene list. Its expression in the cisplatin-resistant group was approximately 50 percent of that in the sensitive group, consistent with prior literature and confirming the biological validity of the broader gene selection framework.

CSP26G also successfully predicted sensitivity to topoisomerase I inhibitors including irinotecan and camptothecin, and to PARP inhibitors including olaparib and talazoparib, but could not predict oxaliplatin sensitivity. This pattern is biologically coherent: oxaliplatin generates DNA adducts through a distinct mechanism, primarily activating nucleolar and ribosomal stress rather than classical mismatch-repair-sensitive DNA lesions, placing it mechanistically outside the spectrum captured by CSP26G.

TL;DR: The 26 CSP26G genes span tumor suppression, apoptosis, drug transport, and immune regulation pathways, and the model extends to predict sensitivity to topoisomerase I and PARP inhibitors while appropriately failing to predict oxaliplatin, which acts through a distinct non-mismatch-repair mechanism.
Pages 7-9
Phenotype-Based Classification as a New Paradigm

The conceptual innovation of CSP26G is phenotype-based classification: rather than identifying a single genetic driver mutation, it captures the integrated gene expression state of the tumor cell that reflects its functional response capabilities across multiple resistance and sensitivity pathways simultaneously. This approach is particularly suited to broad-spectrum cytotoxic agents like cisplatin that engage multiple cellular mechanisms rather than a single targetable pathway.

The integration of SHAP values with DEG analysis represents a methodological advance over prior approaches that used only one analytical framework. SHAP values identify genes whose expression patterns are informationally critical within the multi-gene context of the entire transcriptome, capturing gene-gene interaction effects that single-gene differential expression misses. The intersection of both analyses retains genes that are both statistically different and functionally critical.

The predictive model's clinical value extends beyond predicting cisplatin benefit. Analysis of the cisplatin-resistant group in TCGA revealed significantly lower IC50 values for docetaxel, pemetrexed, and palbociclib, suggesting that CSP26G-predicted resistant patients may benefit from these alternative agents. This points toward a future where the model informs not only cisplatin selection but active selection of alternative first-line therapies for patients unlikely to respond.

Limitations include the retrospective single-cancer-type clinical validation and the requirement for prospective multicenter data with documented cisplatin treatment to establish clinical cutoff values. The current score cutoff range of 0.5 to 0.7 was identified as providing balanced performance in model test data but requires validation in diverse patient cohorts before clinical implementation.

TL;DR: CSP26G establishes phenotype-based transcriptional classification as a viable precision medicine strategy for cytotoxic agents, with SHAP-DEG integration capturing gene interaction effects beyond what single-gene analysis reveals, and resistance predictions pointing toward actionable alternative therapies.
Pages 2, 9, 10
Clinical Translation Potential

The clinical application target for CSP26G is stage IIA to IIIA NSCLC patients who are candidates for cisplatin-based adjuvant chemotherapy after surgical resection, where treatment benefit varies substantially but cannot currently be predicted. In this guideline-defined population, CSP26G stratified TCGA patients into groups with more than a two-year median survival difference, demonstrating clinically meaningful patient-level discrimination.

Implementation requires tumor RNA-seq data, which is increasingly available through standard molecular profiling workflows in oncology centers. Unlike imaging-based or protein-based biomarkers, RNA-seq-derived gene expression scores are readily quantified from surgical or biopsy specimens using widely deployed laboratory platforms.

For patients classified as cisplatin-resistant by CSP26G, the model points toward specific alternative therapeutic options. The observation that resistant-classified cell lines showed lower IC50 values for docetaxel, pemetrexed, and palbociclib provides a data-driven basis for exploring these agents as alternatives, though prospective clinical trials in CSP26G-stratified patients are needed to confirm therapeutic benefit.

Future work will include biological characterization of the 26 biomarker genes to establish mechanistic understanding, application of the same analytical pipeline to other cytotoxic agents, and prospective clinical trials using CSP26G to guide adjuvant chemotherapy selection in resected NSCLC, contributing to the broader goal of bringing precision medicine to classical cytotoxic chemotherapy regimens.

TL;DR: CSP26G targets stage IIA-IIIA NSCLC patients deciding on adjuvant cisplatin therapy, requiring only standard RNA-seq data and providing both a cisplatin response prediction and evidence-based guidance toward alternative agents for predicted non-responders.
Pages 9-10
Toward Precision Medicine for Classical Chemotherapy

This study demonstrates that combining bulk RNA-seq analysis with SHAP-value machine learning can identify gene expression signatures that predict cisplatin sensitivity in a generalizable manner validated across cell lines, experimentally induced resistance models, and clinical outcome data. The CSP26G model achieves a level of predictive performance that supports its development as a clinical biomarker tool.

The phenotype-based approach is fundamentally different from current precision oncology paradigms that target specific driver mutations. By measuring the integrated transcriptional state that reflects a cell's functional capacity for DNA damage response, the model captures information inaccessible to mutation-based genotyping, expanding the scope of actionable biomarker information available to oncologists.

The broader generalizability of CSP26G to topoisomerase I inhibitors and PARP inhibitors, but not to mechanistically distinct platinum agents like oxaliplatin, demonstrates that the model has learned biologically coherent representations of DNA damage response rather than overfitting to cisplatin-specific artifacts. This specificity also provides insight into the mechanistic boundaries of the gene expression signature.

Scaling this approach to additional anticancer agents and cancer types, combined with prospective clinical validation, could establish phenotype-based transcriptomic classifiers as a complement to genomic profiling tests, extending precision medicine principles to the large population of cancer patients who currently receive cytotoxic chemotherapy without predictive biomarker guidance.

TL;DR: CSP26G demonstrates that SHAP-integrated transcriptomic analysis can identify a 26-gene phenotypic signature that predicts cisplatin benefit with validated clinical relevance, extending precision medicine to classical chemotherapy where mutation-based biomarkers currently provide no guidance.
Citation: Open Access, 2025. Available at: PMC12631745.