Identifying Hub Genes in Pancreatic Cancer Using Bioinformatics and Supervised Machine Learning

World J Surg Oncol 2018 AI 6 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Page [1, 2]
Why Hub Gene Discovery Matters in Pancreatic Cancer

Pancreatic ductal adenocarcinoma (PDAC) is driven by complex alterations in gene expression that differ fundamentally between tumor and normal pancreatic tissue. Identifying which genes are most centrally dysregulated can reveal new biomarkers for diagnosis and potential targets for therapeutic intervention.

Not all differentially expressed genes are equally important. Hub genes are highly connected nodes in gene co-expression networks that influence the expression of many other genes. Because of their central network position, hub genes often represent master regulators of disease processes and are more likely to be functionally meaningful than arbitrarily selected differentially expressed genes.

Public gene expression databases such as the Gene Expression Omnibus (GEO) make it possible to study pancreatic cancer gene expression without collecting new patient samples. GEO contains thousands of datasets from cancer research studies worldwide, enabling meta-analyses that span many patient cohorts and increase statistical power.

TL;DR: Hub gene identification using public gene expression data can reveal centrally important regulators of pancreatic cancer biology as candidate biomarkers and drug targets.
Page [2, 3]
Training and Testing Datasets from GEO

The study used three publicly available gene expression datasets from GEO. Datasets GSE16515 and GSE22780 were combined to form the training set, containing a total of 68 samples including both PDAC tumor tissue and matched normal pancreatic tissue.

Dataset GSE15471 served as the independent testing set, comprising 78 samples. Using a completely separate dataset for final validation is critical because it provides an unbiased estimate of how well the identified hub genes and their associated classifiers generalize to new patients.

All three datasets used microarray gene expression profiling technology, which simultaneously measures the expression level of thousands of genes across each patient sample. Preprocessing included background correction, normalization, and log-transformation to ensure comparability of expression values across samples from different platforms.

TL;DR: Two GEO datasets (68 samples total) served as the training set and a third independent GEO dataset (78 samples) served as the testing set for hub gene identification and validation.
Page [3, 4]
Identifying Differentially Expressed Genes and PPI Network Construction

Statistical analysis of gene expression between PDAC tumor samples and normal pancreatic tissue identified 724 differentially expressed genes (DEGs) in the training set. These DEGs were filtered using cutoffs for both fold-change (magnitude of expression difference) and statistical significance (adjusted p-value).

The 724 DEGs were then used to construct a Protein-Protein Interaction (PPI) network using the STRING database, which integrates known and predicted protein interactions from experimental evidence, co-expression data, and literature mining. In the PPI network, each gene is a node and edges represent interactions between their protein products.

Network analysis software was used to calculate connectivity degree for each node in the PPI network, identifying the genes with the highest number of interactions. Genes with exceptionally high degree are the hub nodes that are most deeply embedded in the network and most likely to have broad regulatory influence over other genes.

TL;DR: 724 differentially expressed genes were identified and mapped onto a protein-protein interaction network to find hub nodes with the highest connectivity and regulatory influence.
Pages 6-6
MMP7 and ITGA2 as Pancreatic Cancer Hub Genes

Network analysis identified MMP7 (Matrix Metalloproteinase 7) and ITGA2 (Integrin Alpha 2) as the top hub genes in the pancreatic cancer PPI network. Both genes showed significantly elevated expression in PDAC compared to normal pancreatic tissue across the training datasets.

MMP7 is an enzyme that degrades components of the extracellular matrix, facilitating tumor invasion and metastasis. Its upregulation in PDAC is consistent with the highly invasive nature of pancreatic cancer and its tendency to spread early to surrounding tissues and blood vessels.

ITGA2 encodes a cell surface receptor involved in cell adhesion and migration through its interaction with extracellular matrix proteins. Elevated ITGA2 expression in PDAC contributes to tumor cell motility, invasion into surrounding stroma, and potentially resistance to cell death signals.

TL;DR: MMP7 and ITGA2 were identified as the top hub genes in pancreatic cancer, both mechanistically linked to tumor invasion, metastasis, and interaction with the extracellular matrix.
Pages 8-8
Supervised Machine Learning Validates Hub Gene Classifiers

The study used expression levels of the identified hub genes as features for supervised machine learning classifiers trained to distinguish PDAC from normal tissue. Two algorithms were evaluated: k-Nearest Neighbors (kNN) and Random Forest.

On the independent testing set (GSE15471), the kNN classifier achieved 93.59% accuracy, correctly classifying 73 of 78 samples as either PDAC or normal using only the hub gene expression features. This high accuracy on independent data confirms that MMP7 and ITGA2 expression levels carry strong diagnostic signal.

The Random Forest classifier achieved 81.31% accuracy on the same testing set. The difference in performance between kNN and Random Forest may reflect the small number of features used (hub genes only), for which kNN's distance-based approach may be more appropriate than Random Forest's tree ensemble approach.

The validated diagnostic accuracy of hub gene-based classifiers suggests potential clinical utility as a gene expression biomarker panel for pancreatic cancer diagnosis or risk stratification, particularly if validated using more clinically accessible measurement platforms such as RT-PCR or targeted sequencing.

TL;DR: A kNN classifier using hub gene expression achieved 93.59% accuracy on the independent test set, validating MMP7 and ITGA2 as diagnostically informative biomarkers for pancreatic cancer.
Pages 11-11
Biological Significance and Translational Potential

Both MMP7 and ITGA2 fit established models of PDAC pathobiology. The pancreatic tumor microenvironment is characterized by dense desmoplastic stroma, and proteins involved in extracellular matrix remodeling and cell-matrix adhesion play central roles in enabling PDAC cells to invade and survive in this hostile environment.

From a therapeutic perspective, both hub genes represent potential drug targets. MMP inhibitor drugs have been studied in clinical trials, and integrin-targeting antibodies are an active area of oncology drug development. Elevated MMP7 and ITGA2 expression may also indicate which patients are most likely to respond to therapies targeting these pathways.

A limitation of this study is reliance on microarray data, which has lower dynamic range and sensitivity compared to modern RNA sequencing. Validation using RNA-seq datasets and prospective tissue collections would strengthen confidence in these hub genes as clinically useful biomarkers. Additionally, protein-level validation by immunohistochemistry would confirm that gene expression changes translate into altered protein abundance in tumor tissue.

TL;DR: MMP7 and ITGA2 are biologically plausible hub genes in PDAC and represent promising candidates for diagnostic biomarker panels and therapeutic targeting, pending protein-level validation.
Citation: Open Access, 2018. Available at: PMC6237021.