Gene Prioritization through Consensus Strategy, Enrichment Methodologies Analysis, and Networking for Osteosarcoma Pathogenesis

International Journal of Molecular Sciences 2020 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
The Molecular Puzzle of Osteosarcoma: Why Gene Prioritization Matters

Osteosarcoma (OS) is the most common primary malignant bone tumor, disproportionately affecting adolescents and young adults during peak skeletal growth. Despite decades of research, the molecular etiology of OS has not been pinned down to a discrete set of driver genes. Unlike many other cancer types, OS tumors exhibit extreme chromosomal instability, high somatic mutation rates comparable to breast cancer and leukemia, and complex cytogenetic rearrangements including segment loss, amplification, and translocation, all in the absence of recurrent clonal translocation events. This structural chaos makes it difficult to identify consistent genetic drivers in the way, for instance, BCR-ABL defines CML or KRAS defines a major subset of colorectal cancers.

The heterogeneity problem: OS heterogeneity operates at multiple levels. Individual tumors harbor widely divergent mutational landscapes, and gene expression profiling studies have identified several OS subtypes with distinct biology. High-throughput sequencing efforts have catalogued thousands of mutated genes across OS patient cohorts, but converting this catalogue into a prioritized list of biologically relevant targets requires an integrative computational strategy rather than straightforward frequency counting.

What this study does: Published in the International Journal of Molecular Sciences in 2020, this study from Cabrera-Andrade, Lopez-Cortes, and colleagues applies a consensus gene prioritization strategy that aggregates outputs from nine independent bioinformatics tools, constructs and analyzes a protein-protein interaction (PPI) network from the prioritized genes, and performs communality analysis to identify tightly clustered gene communities. The goal is to move from a broad list of thousands of candidate genes to a compact, biologically coherent set of the most likely OS pathogenic drivers.

The study is fundamentally a systems biology and computational study, not a clinical or wet-lab validation study. Its value lies in organizing existing knowledge into a network framework that reveals how candidate genes relate to one another and to known OS biology, and in identifying a smaller list of high-priority targets for experimental follow-up.

TL;DR: Osteosarcoma is the most common primary bone malignancy, mainly affecting adolescents, and is characterized by extreme genomic instability that has prevented identification of consistent driver genes. This 2020 study uses nine bioinformatics prioritization tools combined with PPI networking and communality analysis to narrow 15,809 candidate genes to 553 consensus genes and then to 47 high-priority targets across six functional gene communities.
Pages 2-4
Nine Tools, One Consensus: How the Prioritization Strategy Was Built

The authors selected nine bioinformatics gene-disease prioritization tools based on two criteria: full availability as web service platforms, and the ability to operate using only the disease name or OMIM code (OMIM 259,500 for OS). The nine methods are Biograph, Cipher, DisGeNET, Genie, GLAD4U, Guildify, Phenolizer, PolySearch, and SNPs3D. Each tool employs a different underlying strategy, ranging from text-mining of biomedical literature (PolySearch, GLAD4U) to network propagation algorithms (Guildify, Cipher) and phenotype-gene mapping (Phenolizer). Importantly, each tool returns a ranked list of genes with an individual score, but the scales and scoring philosophies differ across methods.

Building the consensus score: To integrate these nine independently ranked lists, the authors applied a normalization and geometric mean approach. Each gene's score in each method was normalized, and the final consensus score (Ii) was computed as the geometric mean between the average normalized score across all methods that predicted the gene and a score accounting for how many methods identified the gene. This formulation rewards genes that score consistently across multiple tools over genes that score very highly in just one. The combined output across all nine tools returned 15,809 unique genes, a list too large for meaningful experimental follow-up.

Rational cut-off selection: Rather than applying an arbitrary percentile cut-off, the authors identified the consensus list cut-off by finding the ranking position at which the variation in the Ii score was maximal, representing the sharpest inflection point between genes with high cross-method consensus and those with lower support. This method yielded a cut-off at ranking position 553, reducing the list from 15,809 to 553 genes representing 3.5% of the initial pool.

Validation with curated pathogenic genes: To assess how well the consensus approach captures known biology, the authors manually curated two sets of validated OS pathogenic genes from the literature: G1 genes (47 genes from meta-analyses and case reports in OS patients) and G2 genes (41 genes from animal models and OS cell lines). The consensus method detected 87.2% of G1 genes (41 of 47), 80.5% of G2 genes (33 of 41), and 81.3% of all G1+G2 genes (61 of 75) within the top 553 positions, outperforming every individual tool at the same cut-off. In the top 1% of the consensus list (the top 158 positions), 60% of pathogenic genes (45 of 75) were captured, compared to 35.29% for Genie and 30.14% for Phenolizer, the next best individual methods.

TL;DR: Nine bioinformatics tools (Biograph, Cipher, DisGeNET, Genie, GLAD4U, Guildify, Phenolizer, PolySearch, SNPs3D) were combined using a geometric mean consensus score. An inflection-point cut-off reduced 15,809 candidates to 553 genes. The consensus detected 87.2% of G1 pathogenic genes and 80.5% of G2 genes within those 553 positions, outperforming all nine individual tools. In the top 1%, 60% of all known pathogenic genes appeared, versus 35.3% for the next best individual method.
Pages 4-6
TP53, RB1, and the Top-Ranked Genes in Osteosarcoma

The ten top-ranked genes from the consensus list are TP53, RB1, CHEK2, RUNX2, E2F1, MDM2, CDKN1A, JUN, CCNA2, and CDKN2A. The appearance of TP53 and RB1 at positions 1 and 2 is biologically well-grounded: germline inactivation of both tumor suppressor genes was first described in patients with hereditary retinoblastoma (RB1) and Li-Fraumeni syndrome (TP53), and both were subsequently identified in sporadic OS. As master regulators of the cell cycle, their central role in OS biology has been established over decades of molecular pathology research.

MDM2 and CHEK2: MDM2 ranked 6th. This E3 ubiquitin ligase physically binds to and inactivates TP53, and its gene amplification occurs in 3-25% of primary OS cases, with even higher rates in metastatic and recurrent disease. CHEK2, ranked 3rd, is a DNA damage checkpoint kinase that functions as a TP53 stabilizer and is mutated at a frequency of approximately 7% in OS patients. These two genes underscore a recurring theme in OS pathogenesis: the cell cycle and DNA damage response pathways are under severe selective pressure, and multiple points along these pathways show recurrent disruption.

RUNX2 and osteoblast differentiation: The appearance of RUNX2 at rank 4 reflects the tissue-of-origin biology of OS. RUNX2 is the master transcription factor for osteoblast differentiation, and its deregulation in OS is thought to contribute to the arrested differentiation state of osteosarcoma cells, which retain osteoblast-like characteristics but fail to terminally differentiate into mature bone-forming osteocytes. Loss of normal RUNX2 regulation, combined with TP53 and RB1 inactivation, is a common molecular event in OS.

The mean rank of the 45 pathogenic G1-G2 genes detected in the top 1% of the consensus list was 49.3, meaning these known pathogenic genes were on average located in the top 50 positions out of the full list of 15,809. This tight clustering of biologically validated genes at the top of the ranked list confirms that the consensus strategy successfully concentrates known pathogenic signals toward the head of the prioritization.

TL;DR: Top 10 consensus genes are TP53, RB1, CHEK2, RUNX2, E2F1, MDM2, CDKN1A, JUN, CCNA2, and CDKN2A. MDM2 amplification occurs in 3-25% of primary OS; CHEK2 is mutated in ~7% of patients. Known pathogenic genes averaged rank 49.3 in the top 1% of the 15,809-gene list, confirming signal enrichment at the top of the consensus ranking.
Pages 6-8
PI3K/AKT, MAPK/ERK, and the Signaling Landscape of Osteosarcoma

After establishing the 553-gene consensus list, the authors applied gene ontology (GO) analysis and pathway enrichment analysis using the DAVID Bioinformatics Resource. The GO analysis returned 263 significant biological process terms (FDR-adjusted p-value below 0.01). After filtering with REVIGO to remove redundant terms and retain only those with frequency below 0.01%, 92 non-redundant biological processes remained. The most prominent processes describe DNA replication, cellular proliferation, and apoptotic signaling, with TP53 appearing as the most central signal transducer across these processes. OS-specific biological processes also appeared, including smooth muscle cell and fibroblast proliferation, osteoblast differentiation and development, and positive regulation of mesenchymal cell proliferation, reflecting the mesenchymal tissue-of-origin of osteosarcoma.

KEGG pathway results: Pathway enrichment using the KEGG database placed pathways in cancer, cell cycle, PI3K/AKT, MAPK, FOXO, neurotrophin, and TP53 signaling in the top positions. Reactome enrichment highlighted cell cycle events in more granular detail, including Cyclin D-associated events in G1, Cyclin A/CDK2-associated events at S phase entry, and Cyclin A/B1 events during G2/M transition. Together, these results confirm that the 553-gene consensus list is enriched for genes operating in oncologically relevant signaling circuits rather than being a generic list of highly connected "hub" proteins.

PI3K/AKT and MAPK/ERK deregulation: The Ras/Raf/MEK/ERK pathway is hyperactivated in approximately 30% of all human cancers, and the data presented here show that nearly 67% of OS tumors exhibit aberrant ERK activation, a substantially higher rate than the pan-cancer average. ERK promotes cell proliferation, survival, and metastasis through upstream activation by EGFR and G protein-coupled receptor Ras. Key nodes in this pathway, including SHC1, EGFR, HRAS, PIK3CA, and ERBB2, appear together within Community 10 of the subsequent network analysis, forming a tightly co-expressed regulatory module.

FOXO signaling: The FOXO transcription factor family emerged as a pathway of elevated significance in the k=9 communality analysis compared to the initial consensus enrichment. In normal physiology, PI3K/AKT signaling phosphorylates and inactivates FOXO factors, and gain-of-function PI3K or RAS mutations therefore suppress FOXO activity. In OS specifically, loss of FOXO expression promotes impaired osteogenic differentiation, and FOXO members regulate cell fate via FASLG, TNF apoptosis ligand, and BCL-2 family members (BCL2L1, BNIP3, BCL2L11). The loss of FOXO1 expression in OS tumors is associated with tumor progression and failure of cell cycle arrest.

TL;DR: GO analysis of 553 genes identified 92 non-redundant biological processes (FDR below 0.01), with TP53 as the central signal transducer. KEGG enrichment highlights PI3K/AKT, MAPK/ERK, FOXO, and cell cycle as the top pathways. Aberrant ERK activation is present in approximately 67% of OS tumors, versus 30% in pan-cancer. FOXO signaling emerged as a particularly important downstream target of PI3K/AKT deregulation in the OS-specific network.
Pages 8-10
Building the OS-PPI Network and Measuring Gene Centrality

To move from a flat gene list to a relational network that captures biological interactions, the authors used the STRING database to extract physical protein-protein interactions for human proteins, applying a high-confidence cut-off of 0.9. Of the 553 consensus genes, 505 (91.3%) had at least one qualifying interaction in STRING, forming the OS-PPI network. The 48 genes that did not meet the interaction threshold were excluded from network analysis, a common limitation of PPI network studies that may underrepresent membrane proteins or proteins studied primarily in non-human systems.

Pathogenic genes as network hubs: One of the key analytical steps was comparing the node degree (number of interaction partners) between the 58 known pathogenic genes (G1 and G2 combined) present in the OS-PPI network and the remaining non-pathogenic nodes. Pathogenic genes had a mean node degree of 39.05, compared to 19.25 for non-pathogenic nodes, a difference that was statistically significant by the non-parametric Mann-Whitney U-test (p below 0.001). This result demonstrates that genes validated as pathogenic in OS are more highly connected within the network than the background, supporting the general principle that cancer driver genes tend to occupy hub positions in biological interaction networks.

Top centrality nodes: The centrality index calculation across the OS-PPI network ranked TP53 as the most central node, followed by AKT1, MYC, JUN, EP300, CREBBP, CCND1, CDKN1A, STAT3, and RB1. The convergence of MYC and JUN, two transcription factors with broad roles in proliferation and stress response, alongside TP53, AKT1, and RB1 at the top of the centrality ranking reflects their position as integration points for multiple oncogenic signals rather than simple linear pathway participants.

DRIVE and OncoPPI validation: As a further validation step, the authors compared the OS-PPI network against two independent data resources: the DRIVE project (a large-scale RNAi knockdown screen in 398 cancer cell lines, filtered to 8 bone cancer cell lines including SAOS2, SJSA1, SKES1, and U2OS) and the OncoPPI cancer-focused interaction network. The DRIVE project contained data for 83.5% of the 553 consensus genes (461 of 553). Of these, 20 were classified as essential (dependency score below -3 in more than 50% of bone cancer cell lines), 70 as active, and 371 as inert. The OncoPPI network recognized 92 of the 553 prioritized genes (16.6%), and centrality indices between the OncoPPI and OS-PPI networks showed significant correlation (Spearman r = 0.445, p below 0.001).

TL;DR: The OS-PPI network contains 505 nodes from 553 consensus genes (91.3%). Pathogenic genes average 39.05 interaction partners versus 19.25 for non-pathogenic genes (Mann-Whitney p below 0.001). Top centrality nodes are TP53, AKT1, MYC, JUN, EP300, CREBBP, CCND1, CDKN1A, STAT3, and RB1. DRIVE validation covered 83.5% of consensus genes, with 20 classified essential in bone cancer cell lines. OncoPPI centrality correlation: Spearman r = 0.445, p below 0.001.
Pages 10-12
Six Gene Communities Reveal the Core OS Pathogenic Program

The communality analysis used the clique percolation method implemented in CFinder software to identify overlapping gene communities within the OS-PPI network. Clique percolation detects "k-cliques," groups of fully connected nodes, and then identifies communities as chains of overlapping k-cliques. The method was applied across multiple k values, and the index S (a measure of how evenly genes are distributed across communities) was computed for each. The analysis detected 14 possible k-clique sizes producing 86 possible community configurations with community sizes ranging from 17 to 465 genes.

Selecting k=9: Both k=8 and k=9 produced similar gene distributions (S index of 0.719 and 0.609, respectively), but k=9 was selected because it had a better mean rank of prioritized pathogenic genes (218.89 for k=9 vs. 243.95 for k=8). At k=9, the network yielded 13 communities containing 245 genes (44.3% of the 553-gene consensus list). K-means clustering of these 13 communities based on their mean rank, centrality, and degree scores identified 4 main community groups, with Communities 4, 5, 8, 9, 10, and 13 forming the highest-priority clusters (Clusters 1 and 2).

Composition of the six priority communities: These six communities collectively contain 47 genes, each community having 9 to 13 members with mostly non-overlapping membership. TP53 is the only gene shared across five of the six communities, reinforcing its status as the central regulatory node in OS biology. Communities 4 and 9 are dominated by DNA repair genes: ATM, CHEK1, ATR, BRCA1, BRCA2, RAD51, BLM, and MLH1, members of homologous recombination repair complexes and double-strand break checkpoint pathways. Communities 8, 10, and 13 concentrate PI3K/AKT and MAPK/ERK pathway genes (PIK3CA, PTK2, HRAS, KRAS, SHC1, AKT1) together with the matrix metalloproteases MMP2 and MMP9 and angiogenesis regulators FGF2 and VEGFA. Community 5 contains chromatin remodeling genes ARID1A, SMARCE1, and SMARCB1, pointing to epigenetic dysregulation as a distinct pathogenic axis in OS.

Cross-community validation: From the 47 community genes, 9 were validated by both DRIVE and OncoPPI (BRCA1, AKT1, ATR, CDK4, HRAS, MYC, PIK3CA, RELA, STAT3), 5 by DRIVE alone (RAD51, CDK2, CHEK1, SMARCB1, SMARCE1), and 10 by OncoPPI alone (ATM, CDH1, EGFR, EP300, ERBB2, JUN, NFKB1, SHC1, TP53, SP1). The centrality index of the OS-comms sub-network correlated significantly with the same genes' degree in the full OS-PPI network (r = 0.317, p = 0.03), confirming structural consistency between the sub-network and the full interaction map.

TL;DR: Clique percolation at k=9 yielded 13 communities containing 245 genes. K-means clustering identified 6 priority communities (4, 5, 8, 9, 10, 13) containing 47 genes. TP53 appears in 5 of 6 communities. Communities 4 and 9 are dominated by DNA repair genes (ATM, ATR, CHEK1, BRCA1/2, RAD51, BLM, MLH1). Communities 8, 10, 13 contain PI3K/AKT and MAPK/ERK genes with MMP2/MMP9. Community 5 highlights chromatin remodeling (ARID1A, SMARCB1, SMARCE1). Nine genes were confirmed by both DRIVE and OncoPPI.
Pages 12-14
MMP2/MMP9 in Metastasis and the ATM-ATR-CHEK1 DNA Repair Axis

Pulmonary metastasis is the leading cause of mortality in OS patients, and the molecular events driving metastatic spread are among the most clinically important questions in sarcoma biology. The communality analysis places MMP2 and MMP9 together in Community 13 with a high centrality index. Both matrix metalloproteases degrade extracellular matrix components to facilitate tumor cell migration and invasion, and high MMP9 expression has been observed specifically in metastatic OS samples. Their co-localization with upstream regulators of the MAPK/ERK cascade (IL6, FGF2, VEGFA, EGFR, ERBB2) in Community 13 suggests these proteases are downstream effectors of PI3K/AKT and MAPK/ERK signaling rather than independently regulated factors.

The metastatic cascade: The full metastatic program in OS involves cellular detachment from the primary tumor, matrix remodeling and invasion, angiogenesis, vascular dissemination, and proliferation at distant sites. The gene composition of Community 13 addresses multiple steps in this cascade simultaneously: VEGFA supports angiogenesis, MMP2 and MMP9 handle matrix remodeling, and EGFR/ERBB2 sustain proliferative signaling at secondary sites. FGF2 and IL6, which appear as upstream MAPK/ERK regulators in the prioritization, have established roles in promoting both angiogenesis and immune evasion in sarcomas.

Homologous recombination and the ALT pathway: Communities 4 and 9 highlight a second major pathogenic axis: the DNA damage response and homologous recombination (HR) repair complex. ATM-CHEK2 controls cellular responses to double-strand breaks, while ATR-CHEK1 responds to DNA replication stress via phosphorylation cascades triggered by UV damage and replication fork stalling. In the OS-comms network, ATM, ATR, and CHEK1 all have high centrality indices and interact with BRCA1 and RAD51 (classified as essential by DRIVE) and with CDK2 and CDK4 (classified as active).

Alternative lengthening of telomeres (ALT): Osteosarcoma is one of the cancer types most commonly associated with the ALT pathway for telomere maintenance, used in place of telomerase activation. ALT relies on HR-based mechanisms, and several proteins localized at ALT-associated promyelocytic leukemia bodies (APBs) appear prominently in the prioritization: PML, the RecQ family helicases BLM, WRN, and RECQL4, RAD51, and RAD52. The authors argue that the HR repair complex serves dual pathogenic roles in OS: it contributes to cell cycle disruption through checkpoint deregulation, and it enables the chromosome instability and alternative telomere biology that are hallmarks of this tumor type specifically.

TL;DR: Community 13 places MMP2 and MMP9 within the PI3K/AKT and MAPK/ERK network alongside VEGFA, EGFR, ERBB2, FGF2, and IL6, linking matrix remodeling to upstream kinase deregulation. High MMP9 expression correlates with metastatic OS samples. Communities 4 and 9 highlight the ATM-ATR-CHEK1 DNA repair axis and its interaction with BRCA1/2 and RAD51. ALT-associated proteins (BLM, WRN, RECQL4, RAD51, RAD52) appear prominently, connecting HR repair biology to OS-specific telomere maintenance.
Pages 14-16
Transcription Factor Analysis, Key Limitations, and Future Directions

The authors recognized that PPI-based prioritization can be biased toward proteins with many physical interactions, potentially underweighting transcription factors (TFs) whose influence operates primarily through gene regulatory mechanisms rather than direct protein binding. To address this, they performed a secondary prioritization focused exclusively on the 125 TFs identified within the 553-gene consensus list, using ChIP-seq binding site data from the GTRD (Gene Transcription Regulation Database) to score each TF by how many of the 553 consensus genes it regulates. More than 82% of identified TFs (103 of 125) were found to regulate more than half of all consensus genes simultaneously, indicating pervasive transcriptional co-regulation across the prioritized gene set.

Top transcription factors: The TF re-ranking placed TP53, E2F1, JUN, RUNX2, FLI1, YY1, HIF1A, MYC, TP63, ESR1, WT1, E2F4, ATF2, NFKB1, AR, SP1, STAT1, ERG, CEBPB, and TFAP2A in the top 20 positions. Notably, E2F1 and E2F4 improved substantially in ranking compared to the main consensus list, consistent with their role as direct RB1 transcriptional targets. During G1, RB1 suppresses E2F activity; cyclin-dependent kinase phosphorylation of RB1 by CDK4/CDK6 and CDK2 releases E2F transcription factors to drive expression of cell cycle genes including Cyclins A, D, and E. The improved score of these TFs suggests that E2F-mediated cell cycle deregulation is a foundational event in OS pathogenesis downstream of RB1 loss. NFkB1 also emerges as a central regulatory node in Communities 5, 8, and 13, reflecting its known roles in promoting cell survival, autophagy, and osteoblast differentiation via interactions with TGFB1 and BCL2 family members.

Limitations of the study: The study is entirely computational, and all findings require experimental validation. The consensus prioritization strategy depends on the completeness and accuracy of the nine input tools, which in turn depend on the biomedical literature and databases available at the time of study. Genes that are biologically important but poorly studied (and therefore underrepresented in databases) would be systematically deprioritized. PPI data from STRING reflects interactions annotated in the literature, introducing a similar "rich get richer" bias toward well-studied proteins. The DRIVE project data covers eight bone cancer cell lines, which may not fully represent the in vivo heterogeneity of patient tumors. Additionally, this study does not account for racial, ethnic, or age-related genetic factors that may influence OS biology and prevalence.

Future directions: The authors recommend experimental validation of prioritized genes, particularly those in Communities 4 and 9 (ATM, ATR, CHEK1, RAD51, BLM) and the chromatin remodeling genes of Community 5 (ARID1A, SMARCB1, SMARCE1), which have received less attention than the PI3K/AKT pathway genes. Investigation of genetic variants in the identified transcription factors and their relationship to OS incidence across demographic groups is also proposed as an important area. The ALT pathway connection, linking HR repair complex members to telomere maintenance biology specific to bone tumors, represents a particularly interesting therapeutic angle given the emergence of PARP inhibitors and ATR inhibitors in clinical trials for HR-deficient cancers.

TL;DR: Secondary TF prioritization using GTRD ChIP-seq data ranked TP53, E2F1, JUN, RUNX2, NFKB1, and MYC as top regulators; 82.4% of identified TFs regulate more than half of all consensus genes. Key limitations include full reliance on existing databases (literature bias), PPI "hub" bias, and cell-line-based DRIVE data that may not represent patient tumor heterogeneity. Priority experimental targets are ATM/ATR/CHEK1 (DNA repair), ARID1A/SMARCB1/SMARCE1 (chromatin remodeling), and ALT pathway components as potential therapeutic vulnerabilities.