Development and validation of a 16-gene T-cell-related prognostic model in non-small cell lung cancer

Front Immunol 2025 AI 6 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Page 1
A 16-Gene Immune Signature to Predict Survival and Guide Treatment in NSCLC

The problem with current NSCLC prognosis tools Non-small cell lung cancer has a 5-year survival rate of only ~24% overall and ~6% in advanced stages. While molecular markers (gene mutations, protein expression) exist, they often fail to capture the critical role of T-cell activity in the tumor immune microenvironment - a major driver of how patients respond to immunotherapy and chemotherapy.

Study approach Researchers analyzed RNA sequencing data from 1,027 NSCLC and 108 non-cancerous tissue samples from The Cancer Genome Atlas (TCGA). Using weighted gene co-expression network analysis (WGCNA), immune cell scoring (ssGSEA), and LASSO Cox regression, they identified a 16-gene signature specifically linked to T-cell activity and built it into a prognostic risk model.

The 16-gene model The final prognostic model includes: LATS2, LDHA, CKAP4, COBL, DSG2, MAPK4, AKAP12, HLF, CD69, BAIAP2L2, FSTL3, CXCL13, PTX3, SMO, KREMEN2, and HOXC10. These genes were selected for their strong association with T-cell activity and NSCLC outcomes.

Key results The model stratified patients into high-risk and low-risk groups with significant survival differences. Time-dependent AUCs were 0.68, 0.72, and 0.69 for 1-, 3-, and 5-year survival prediction. External validation on three GEO datasets confirmed the model's robustness. Drug sensitivity analysis suggested high-risk patients may benefit from MCL1 inhibitors, while low-risk patients may respond better to IGF1R inhibitors.

TL;DR: A bioinformatics pipeline identified a 16-gene T-cell-related prognostic model for NSCLC that stratifies patients into high and low risk groups, predicts 1-, 3-, and 5-year survival, and suggests tailored drug treatment strategies for each risk group.
Page 2
Why T-Cell Activity is Critical for NSCLC Prognosis and Immunotherapy Response

T cells as key survival determinants The immune system's ability to recognize and destroy cancer cells is largely driven by cytotoxic CD8+ T cells and helper CD4+ T cells. In NSCLC, high T-cell infiltration correlates with better prognosis and stronger response to immune checkpoint inhibitors (PD-1/PD-L1 and CTLA-4 inhibitors). However, T-cell function is highly variable between patients.

Limitations of existing prognostic tools Current NSCLC molecular classifiers based on driver gene mutations (EGFR, KRAS, ALK) or standard histology do not fully reflect the tumor immune microenvironment. Immunohistochemical T-cell assessments exist, but they are qualitative, subject to observer variability, and do not capture the complexity of T-cell-related gene expression patterns.

Gap that this study fills Transcriptomic profiling of T-cell-related genes provides a more comprehensive and quantifiable view of immune activity. A prognostic model built from T-cell-associated gene expression could complement or improve upon existing staging and molecular biomarkers - particularly for guiding immunotherapy decisions.

Key analytical tools used ssGSEA scored immune cell abundance per patient; WGCNA identified clusters of co-expressed genes correlated with T-cell activity; LASSO Cox regression then selected the most prognostically relevant genes from hundreds of candidates, penalizing model complexity to prevent overfitting.

TL;DR: T-cell activity in the NSCLC tumor microenvironment is a key driver of prognosis and immunotherapy response, but existing molecular classifiers fail to fully capture this, motivating the development of a T-cell gene expression-based prognostic model.
Pages 2-4
Building the Prognostic Model: From TCGA Data to LASSO Cox Regression

Data and cohorts TCGA RNA sequencing data from 1,027 NSCLC cancer samples and 108 non-cancerous samples were the primary training dataset. Three external GEO datasets (GSE50081, n=173; GSE31210, n=225; GSE30219, n=190) provided independent validation cohorts. Data were batch-corrected using the ComBat function to remove technical variability between datasets.

Gene discovery pipeline ssGSEA first scored 28 immune cell subtype abundances per sample. WGCNA then identified gene co-expression modules correlated with T-cell abundance (correlation at least 0.5). Hub genes in these modules were identified by high gene significance and module membership scores. Differentially expressed genes between NSCLC and normal tissue were intersected with T-cell hub genes to yield 80 T-cell-related DEGs as candidates.

LASSO Cox model construction The TCGA cohort was split 7:3 for training/testing. LASSO Cox regression applied L1 regularization to select the most prognostically informative genes from the 80 candidates, penalizing model complexity to reduce overfitting. The optimal penalty parameter was selected by cross-validation. Sixteen genes with non-zero coefficients formed the final risk score formula.

Validation and downstream analyses The 16-gene risk score was validated on three external GEO cohorts. A nomogram combining the risk score with clinical variables (age, stage, smoking) was constructed. Drug sensitivity was predicted using the oncoPredict package (GDSC2 and CTRP databases). Gene expression of model genes was confirmed in 8 NSCLC tissue samples by quantitative RT-PCR.

TL;DR: The 16-gene model was built through a systematic ssGSEA to WGCNA to LASSO Cox regression pipeline on 1,027 TCGA NSCLC samples and validated on three independent GEO datasets, with qRT-PCR confirmation in clinical tissue samples.
Pages 5-7
Risk Stratification, Survival Prediction, and Immune Microenvironment Differences

Survival stratification Using the median risk score as the cutoff, patients were divided into high-risk and low-risk groups with significantly different overall survival (log-rank p less than 0.05) in both the TCGA training/test cohorts and all three GEO validation datasets. High-risk patients consistently had worse survival outcomes across all datasets, confirming the model's prognostic utility.

Time-dependent AUC performance The model achieved AUCs of 0.68, 0.72, and 0.69 for 1-, 3-, and 5-year overall survival prediction in the training cohort. The nomogram combining risk score with clinical factors (age, stage) improved prediction (AUC greater than 0.6 in all validation cohorts), though the improvement over the gene score alone was modest.

Immune microenvironment differences High-risk patients had distinct immune cell composition: lower levels of activated CD8+ T cells and higher levels of immunosuppressive cells such as regulatory T cells and M2 macrophages. GSEA identified enrichment of proliferation and metabolism-related Hallmark pathways in high-risk patients, consistent with more aggressive tumor biology.

Drug sensitivity findings High-risk patients showed predicted sensitivity to AZD5991-1720, an MCL1 inhibitor that induces apoptosis in cancer cells with anti-apoptotic dependencies. Low-risk patients showed predicted sensitivity to IGF1R-3801-1738, an IGF1R inhibitor. This suggests the risk model could guide selection of non-immunotherapy drugs tailored to each patient's tumor biology.

TL;DR: The 16-gene model stratified NSCLC patients into distinct survival groups across four independent cohorts, with high-risk tumors showing immunosuppressive microenvironments and predicted sensitivity to different drugs than low-risk tumors.
Pages 7-8
What the 16 Genes Reveal About T-Cell Biology in NSCLC

Immune activation genes CD69 is an early T-cell activation marker, while CXCL13 is a chemokine that attracts B cells and promotes tertiary lymphoid structure formation - both associated with active antitumor immunity. PTX3, a regulator of innate immunity, plays roles in complement activation. HLF is linked to hematopoietic stem cell maintenance and immune regulation.

Metabolic and structural regulators LDHA (lactate dehydrogenase A) drives the Warburg effect - the shift of cancer cells to aerobic glycolysis - creating an acidic microenvironment hostile to T-cell function. LATS2 is a tumor suppressor in the Hippo signaling pathway. CKAP4 is a cytoskeleton-associated protein with roles in cell signaling and migration.

Immunosuppression and evasion genes FSTL3 is a TGF-beta superfamily antagonist that modulates immune responses. MAPK4 is involved in MAPK signaling relevant to T-cell activation and tumor proliferation. HOXC10 is a transcription factor implicated in cancer progression and therapy resistance. DSG2 encodes desmoglein 2, a component of desmosomes involved in epithelial integrity and tumor cell dissemination.

Pathway enrichment The set of 16 genes was enriched in pathways related to T-cell activation, immune response regulation, and tumor metabolic reprogramming - collectively supporting the concept that this signature captures the interface between T-cell biology and tumor immune evasion mechanisms in NSCLC.

TL;DR: The 16 model genes collectively capture key aspects of T-cell activation (CD69, CXCL13), tumor metabolic immunosuppression (LDHA), tumor suppressor signaling (LATS2), and immune regulatory pathways that together determine the tumor immune microenvironment phenotype.
Pages 8-9
Study Constraints and the Path Toward Clinical Validation

Retrospective public database design The model was built entirely from publicly available TCGA and GEO datasets. These databases have inherent limitations including variable follow-up lengths, missing clinical data (treatment details, recurrence information), and batch effects between studies that may not be fully corrected even with ComBat normalization.

Limited tissue validation The qRT-PCR validation included only 8 lung adenocarcinoma tissue samples - far too small to draw clinical conclusions. This tissue validation serves as a preliminary confirmation of gene expression trends rather than a clinical validation of the prognostic model itself.

AUC performance limitations Time-dependent AUCs of 0.68-0.72 indicate moderate predictive accuracy. While better than chance, these values fall short of the 0.80+ threshold typically considered clinically robust. The model's practical advantage over established staging (TNM) and existing molecular markers needs to be tested in head-to-head comparisons with clinical data.

Future directions Prospective multicenter clinical trials are essential to validate the 16-gene model in real-world NSCLC patients with standardized treatment protocols. Integration with existing clinical variables and molecular biomarkers (PD-L1, TMB, driver mutations) could create a composite model with greater predictive power. Translation to a validated assay platform for clinical use requires regulatory pathway development.

TL;DR: The model's reliance on retrospective public datasets, limited tissue validation, and moderate AUC values (0.68-0.72) require future prospective multicenter studies and head-to-head comparisons with standard prognostic tools before any clinical adoption.
Citation: Open Access, 2025. Available at: PMC12009871.