Pancreatic cancer lacks reliable biomarkers — measurable biological signals that indicate cancer is present. Identifying specific genes whose expression is abnormal in pancreatic cancer tissue could provide both diagnostic markers and drug targets.
Risk factors like smoking, obesity, diabetes, and excessive alcohol use are known to increase pancreatic cancer risk, but the molecular mechanisms linking these factors to cancer development are poorly understood.
This study used a neural network model applied to publicly available gene expression datasets to identify 'feature genes' — genes whose expression levels most strongly and consistently differ between pancreatic cancer and normal pancreatic tissue.
Transcriptome (gene expression) data were downloaded from three public GEO datasets: GSE15471, GSE16515 (combined as training group), and GSE32676 (independent test group). Survival data for 178 pancreatic cancer patients from TCGA were also analyzed.
Differentially expressed genes (DEGs) between cancer and normal tissue were identified first. The neural network then scored each DEG by importance, selecting the top 10 feature genes with importance scores above 2.
The neural network used gene importance scores as inputs, built hidden layers from these weights, and produced a binary output classifying samples as cancer or normal. ROC curves assessed classification accuracy.
The neural network identified 10 feature genes with the highest discriminatory power: FGD6, ANO1, POSTN, AHNAK2, FN1, SLC39A5, RHBDL2, MTMR11, SQLE, and ADAM9. These genes were validated at both the transcriptome (RNA) and protein levels.
GO enrichment analysis revealed these genes are involved in epidermis development, endoplasmic reticulum function, and sulfur compound binding. KEGG pathway analysis linked them to pancreatic secretion and complement/coagulation cascades — processes relevant to tumor biology and immune evasion.
Protein-protein interaction (PPI) network analysis showed many of these feature genes are closely connected, suggesting they form a functionally related gene module rather than acting independently — a sign that disruption of this network may drive cancer progression.
Survival analysis using TCGA data showed that expression levels of the feature genes were significantly associated with patient prognosis. Patients with higher expression of certain genes had shorter survival, confirming these genes are prognostically relevant.
Immune cell infiltration analysis using the CIBERSORT algorithm revealed that feature genes correlate with the composition of immune cells in the tumor microenvironment. Notably, neutrophil activity increased while other immune populations shifted in tumors with abnormal feature gene expression.
A correlation heatmap of immune cell types showed that altered feature gene expression reshapes the entire immune landscape of the tumor, potentially contributing to immune evasion and treatment resistance.
Neural network feature selection on public gene expression datasets can efficiently identify a small, highly informative set of genes that captures the biology of pancreatic cancer. The 10-gene panel shows promise as both a diagnostic and prognostic tool.
Because feature genes were validated across multiple independent datasets and at the protein level, findings are more robust than studies relying on a single dataset. Confirmation of prognostic relevance in TCGA survival data adds further clinical credibility.
Future work should test whether blood-based or biopsy-based measurement of these feature genes can be developed into a practical clinical assay for early pancreatic cancer detection or treatment monitoring.