Pancreatic ductal adenocarcinoma (PDAC) is driven by complex alterations in gene expression that differ fundamentally between tumor and normal pancreatic tissue. Identifying which genes are most centrally dysregulated can reveal new biomarkers for diagnosis and potential targets for therapeutic intervention.
Not all differentially expressed genes are equally important. Hub genes are highly connected nodes in gene co-expression networks that influence the expression of many other genes. Because of their central network position, hub genes often represent master regulators of disease processes and are more likely to be functionally meaningful than arbitrarily selected differentially expressed genes.
Public gene expression databases such as the Gene Expression Omnibus (GEO) make it possible to study pancreatic cancer gene expression without collecting new patient samples. GEO contains thousands of datasets from cancer research studies worldwide, enabling meta-analyses that span many patient cohorts and increase statistical power.
The study used three publicly available gene expression datasets from GEO. Datasets GSE16515 and GSE22780 were combined to form the training set, containing a total of 68 samples including both PDAC tumor tissue and matched normal pancreatic tissue.
Dataset GSE15471 served as the independent testing set, comprising 78 samples. Using a completely separate dataset for final validation is critical because it provides an unbiased estimate of how well the identified hub genes and their associated classifiers generalize to new patients.
All three datasets used microarray gene expression profiling technology, which simultaneously measures the expression level of thousands of genes across each patient sample. Preprocessing included background correction, normalization, and log-transformation to ensure comparability of expression values across samples from different platforms.
Statistical analysis of gene expression between PDAC tumor samples and normal pancreatic tissue identified 724 differentially expressed genes (DEGs) in the training set. These DEGs were filtered using cutoffs for both fold-change (magnitude of expression difference) and statistical significance (adjusted p-value).
The 724 DEGs were then used to construct a Protein-Protein Interaction (PPI) network using the STRING database, which integrates known and predicted protein interactions from experimental evidence, co-expression data, and literature mining. In the PPI network, each gene is a node and edges represent interactions between their protein products.
Network analysis software was used to calculate connectivity degree for each node in the PPI network, identifying the genes with the highest number of interactions. Genes with exceptionally high degree are the hub nodes that are most deeply embedded in the network and most likely to have broad regulatory influence over other genes.
Network analysis identified MMP7 (Matrix Metalloproteinase 7) and ITGA2 (Integrin Alpha 2) as the top hub genes in the pancreatic cancer PPI network. Both genes showed significantly elevated expression in PDAC compared to normal pancreatic tissue across the training datasets.
MMP7 is an enzyme that degrades components of the extracellular matrix, facilitating tumor invasion and metastasis. Its upregulation in PDAC is consistent with the highly invasive nature of pancreatic cancer and its tendency to spread early to surrounding tissues and blood vessels.
ITGA2 encodes a cell surface receptor involved in cell adhesion and migration through its interaction with extracellular matrix proteins. Elevated ITGA2 expression in PDAC contributes to tumor cell motility, invasion into surrounding stroma, and potentially resistance to cell death signals.
The study used expression levels of the identified hub genes as features for supervised machine learning classifiers trained to distinguish PDAC from normal tissue. Two algorithms were evaluated: k-Nearest Neighbors (kNN) and Random Forest.
On the independent testing set (GSE15471), the kNN classifier achieved 93.59% accuracy, correctly classifying 73 of 78 samples as either PDAC or normal using only the hub gene expression features. This high accuracy on independent data confirms that MMP7 and ITGA2 expression levels carry strong diagnostic signal.
The Random Forest classifier achieved 81.31% accuracy on the same testing set. The difference in performance between kNN and Random Forest may reflect the small number of features used (hub genes only), for which kNN's distance-based approach may be more appropriate than Random Forest's tree ensemble approach.
The validated diagnostic accuracy of hub gene-based classifiers suggests potential clinical utility as a gene expression biomarker panel for pancreatic cancer diagnosis or risk stratification, particularly if validated using more clinically accessible measurement platforms such as RT-PCR or targeted sequencing.
Both MMP7 and ITGA2 fit established models of PDAC pathobiology. The pancreatic tumor microenvironment is characterized by dense desmoplastic stroma, and proteins involved in extracellular matrix remodeling and cell-matrix adhesion play central roles in enabling PDAC cells to invade and survive in this hostile environment.
From a therapeutic perspective, both hub genes represent potential drug targets. MMP inhibitor drugs have been studied in clinical trials, and integrin-targeting antibodies are an active area of oncology drug development. Elevated MMP7 and ITGA2 expression may also indicate which patients are most likely to respond to therapies targeting these pathways.
A limitation of this study is reliance on microarray data, which has lower dynamic range and sensitivity compared to modern RNA sequencing. Validation using RNA-seq datasets and prospective tissue collections would strengthen confidence in these hub genes as clinically useful biomarkers. Additionally, protein-level validation by immunohistochemistry would confirm that gene expression changes translate into altered protein abundance in tumor tissue.