Identification and Verification of Immune and Oxidative Stress-Related Diagnostic Indicators for Malignant Lung Nodules Through WGCNA and Machine Learning

Sci Rep 2025 AI 5 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Finding the Molecular Fingerprint That Separates Malignant From Benign Lung Nodules

The Clinical Problem With Lung Nodules Lung nodules detected on CT scans are common but diagnostically ambiguous. Experienced radiologists still struggle to reliably distinguish benign from malignant lesions, and subjective interpretation variability leads to both unnecessary interventions and missed cancers. There is an urgent unmet need for molecular biomarkers that can objectively identify which nodules are truly malignant.

Why Immunity and Oxidative Stress? Two biological pathways are increasingly recognized as central to lung cancer development. Oxidative stress - the imbalance between reactive oxygen species (ROS) production and antioxidant defenses - causes genomic instability and aberrant cell growth that drive malignant transformation. Simultaneously, the immune system's ability to recognize and eliminate cancer cells is dysregulated within the lung nodule microenvironment, with distinct immune cell infiltration patterns distinguishing malignant from benign lesions.

WGCNA as a Systems-Level Tool Weighted gene co-expression network analysis (WGCNA) identifies groups of genes that are co-regulated under specific biological conditions. Rather than examining individual genes in isolation, WGCNA reveals entire modules of functionally related genes that change together in malignant versus benign lung nodules. This systems-level view captures the coordinated biology of malignant transformation rather than single-gene effects.

113 Machine Learning Combinations Rather than selecting a single machine learning algorithm based on convention, this study systematically tested 113 combinations of 12 different algorithms to identify the optimal diagnostic model. Only the combination with the highest area under the curve (AUC) in both training and external validation datasets was selected, providing confidence that the final model was not an artifact of algorithm choice.

TL;DR: This study combined WGCNA and 113 machine learning algorithm combinations to identify an 11-gene immune and oxidative stress-related diagnostic signature for distinguishing malignant from benign lung nodules, outperforming existing published models.
Pages 2-4
Layered Bioinformatics: From Gene Expression to Diagnostic Model

Datasets and Batch Effect Correction GSE108375 (164 malignant, 151 benign nodule samples) served as the training dataset, and GSE20189 (81 malignant, 18 benign) served as external validation. Both datasets were merged and batch effects corrected using the ComBat function from the sva R package, confirmed by principal component analysis showing pre- and post-correction separation. Differentially expressed genes (DEGs) were identified using the limma package with thresholds of |log2 fold-change| > 0.1 and FDR < 0.05, yielding 1,432 DEGs (857 upregulated, 575 downregulated).

Intersecting Immunity and Oxidative Stress Genes Oxidative stress-related genes were obtained from the GeneCards database using a relevance score threshold of 7 or above, yielding 817 candidate genes. WGCNA was independently run twice: once correlating gene modules with the clinical phenotype (malignant vs. benign), and once correlating modules with immune cell infiltration profiles quantified by ssGSEA. The modules most strongly associated with each phenotype were intersected with the 1,432 DEGs and with the 817 oxidative stress genes. This three-way intersection yielded 31 differentially expressed immune and oxidative stress-related genes (DEIOGs).

PPI Network and Hub Gene Identification A protein-protein interaction (PPI) network was constructed from the 31 DEIOGs using the STRING database. Hub genes were identified by applying all 12 centrality-scoring algorithms available in the CytoHubba plugin within Cytoscape - including degree, betweenness, closeness centrality, and others. Only genes that ranked in the top 10 across multiple algorithms were retained as consensus hub genes, yielding CDK2 and MCL1 as the two hub genes.

Machine Learning Model Development Using the 31 DEIOGs as input features, 12 machine learning algorithms were tested: naive Bayes, SVM, glmBoost, LDA, ridge regression, LASSO, stepglm, plsRglm, GBM, elastic net, random forest, and XGBoost. All 113 pairwise and combination configurations were evaluated using 10-fold cross-validation on the 70% training partition of GSE108375, with the remaining 30% as a test set and GSE20189 as external validation. The Stepglm[both] + Random Forest combination achieved the highest AUC across all three evaluation sets.

TL;DR: After batch effect correction and limma-based DEG identification, a three-way intersection of DEGs, WGCNA modules, and 817 oxidative stress genes produced 31 DEIOGs; 113 ML combinations were tested to select Stepglm+RF as the optimal diagnostic model.
Pages 5-7
An 11-Gene Signature and Two Hub Genes With Opposing Risk Associations

The 11-Gene Diagnostic Signature The Stepglm[both] + Random Forest model selected 11 genes from the 31 DEIOGs: IL1R1, GADD45A, SMAD2, AHSP, ENDOG, CLEC4A, SHC1, SDHAF1, PXN, SUOX, and BAD. Seven of these (SDHAF1, GADD45A, AHSP, CLEC4A, SHC1, SUOX, and IL1R1) were upregulated in malignant nodules, while ENDOG, BAD, and SMAD2 were downregulated. BAD expression additionally varied significantly across pathological stages, linking it to malignancy progression severity.

Outperforming Existing Models The 11-gene model demonstrated superior diagnostic performance compared to the previously published Guan's XGBoost model and Nishio's gradient tree boosting model in both training and external validation cohorts. DeLong test comparisons confirmed statistically significant AUC improvements (P < 0.05). Calibration curves showed strong agreement between predicted probabilities and observed outcomes, confirming both discriminative and calibration quality.

Immune Cell Infiltration Patterns ssGSEA analysis of the GSE108375 dataset revealed that malignant nodules showed elevated proportions of mast cells, natural killer T cells, neutrophils, and plasmacytoid dendritic cells compared to benign controls. Conversely, activated B cells, activated CD4+ T cells, eosinophils, and natural killer cells were reduced in malignant nodules. These distinct immune infiltration profiles support the hypothesis that immune dysregulation is a hallmark of nodule malignancy.

CDK2 and MCL1 Hub Gene Clinical Correlations The two hub genes showed opposing clinical associations with pack-years (cumulative smoking exposure). CDK2 overexpression correlated with longer pack-year history (P = 0.012), possibly reflecting its role in DNA repair adaptation to smoking-induced genotoxic damage. MCL1 overexpression was associated with lower pack-years (P = 0.018), consistent with its anti-apoptotic function promoting survival of minimally smoke-damaged but oncogenically transformed cells. In nodule size stratification analysis (greater than or equal to 10 mm as high-risk), CDK2 clustered with the low-risk group while MCL1 clustered with the high-risk group.

TL;DR: The Stepglm+RF 11-gene model outperformed Guan's and Nishio's published models in both training and validation; CDK2 and MCL1 hub genes showed opposing associations with smoking exposure and nodule risk stratification.
Pages 8-10
Biological Roles of the 11 Signature Genes in Lung Cancer

Immune and Checkpoint-Related Genes CLEC4A is differentially expressed in peripheral blood mononuclear cells of early-stage lung cancer patients, suggesting blood-detectable utility. IL1R1, encoding the interleukin-1 receptor, showed higher T-cell expression in lung cancer biopsies than in pleural effusions by single-cell RNA sequencing. PXN (paxillin) is a downstream mediator of CXCL5 signaling that activates AKT to upregulate PD-L1 expression in lung cancer cells, directly linking it to immune checkpoint evasion and CD8+ T cell exhaustion.

Apoptosis and Cell Cycle Regulators GADD45A is upregulated by specific agents in lung cancer cells and is associated with cell cycle arrest and apoptosis, making its overexpression in malignant nodules paradoxically linked to adaptive cancer cell responses. BAD downregulation may enhance survival signaling in the EGFR-TKI treatment context. SMAD2, involved in TGF-beta signaling, has been identified as a therapeutic target for metastatic lung cancer through the GATM/Smad2 pathway.

DNA Damage and Mitochondrial Genes ENDOG plays a role in the DNA damage response (DDR) pathway, particularly in tumor repopulation after radiotherapy through the Cox-2/PGE2 axis. Its downregulation in malignant nodules may disrupt pro-tumorigenic DDR mechanisms. SHC1 is implicated in crizotinib resistance in ALK-driven lung cancer, and SUOX has been proposed as a molecular target for lung cancer therapy. SDHAF1 and AHSP were the two genes in the signature without established lung cancer literature associations, flagged as requiring experimental validation.

Pathway-Level Context From Enrichment Analysis GO and KEGG enrichment of the 31 DEIOGs highlighted pathways including positive regulation of mitochondrial organization, autophagy, cytokine signaling, lipid metabolism, and carbohydrate metabolism. GSEA identified enrichment for the CDC42 GTPase cycle and Fc-gamma receptor-dependent phagocytosis, connecting the gene set to cell motility and innate immune phagocytic signaling. Disease ontology analysis linked the DEIOGs to mitochondrial metabolism disorders and oxidative phosphorylation deficiency - pathways relevant to both cancer progression and treatment resistance.

TL;DR: The 11 signature genes span immune checkpoint regulation (PXN, IL1R1, CLEC4A), apoptosis and cell cycle (BAD, GADD45A, SMAD2), DNA damage response (ENDOG, SHC1), and mitochondrial/metabolic function (SDHAF1, SUOX, AHSP).
Pages 11-12
From Bioinformatics to Clinically Validated Molecular Testing

Reliance on Public Datasets The study relied entirely on publicly available GEO datasets. GSE108375 (315 samples) and GSE20189 (99 samples) are relatively small for deep learning standards, and the inherent biases of retrospectively collected public data - including variable sample collection protocols, platform differences, and population heterogeneity - limit the reliability of generalized conclusions. Prospective multicenter cohort validation with standardized sample collection is essential.

Two Genes Lacking Literature Precedent AHSP and SDHAF1, two of the 11 machine learning-selected genes, have no established associations with lung cancer in the published literature. While their selection by both stepglm and random forest algorithms provides algorithmic support, they require experimental validation using qPCR or western blotting in clinical lung cancer samples before their biological roles can be interpreted or trusted.

No Functional Validation Experiments Unlike many biomarker studies, this work did not include in vitro knockdown or knockout experiments to confirm that the identified genes causally drive malignant behavior. CRISPR-Cas9-mediated knockout of CDK2, MCL1, and other hub genes in lung cancer cell lines could clarify their functional roles. Animal model experiments would further confirm whether modulating these genes alters tumor growth or progression.

Therapeutic and Multi-Omics Extensions The identified genes span pathways already under active therapeutic investigation - CDK2 inhibitors for metastatic lung cancer, MCL1 inhibitors for overcoming EGFR-TKI resistance. Future work should determine whether expression of these 11 genes predicts patient response to immunotherapy or oxidative stress-modulating agents. Integration with proteomics, metabolomics, or epigenomics data may reveal additional diagnostic or therapeutic co-targets within the same pathways.

TL;DR: Prospective multicenter validation, experimental confirmation of AHSP and SDHAF1, functional knockdown studies, and integration with therapeutic response data are the critical next steps for translating this bioinformatics framework into clinical lung nodule diagnostics.
Citation: Open Access, 2025. Available at: PMC12218290.