Cancer remains the second leading cause of death worldwide, with nearly 9.7 million deaths recorded in 2022 alone. One of the most pressing diagnostic challenges is accurately identifying the tissue of origin (TOO) of a tumor -- that is, determining which organ a cancer started in -- particularly when tumors have spread or present with non-specific symptoms.
MicroRNAs (miRNAs) are small molecules, typically 17 to 25 nucleotides long, that regulate gene expression throughout the body. They have emerged as powerful cancer biomarkers because they can act as either oncogenes (promoting cancer) or tumor suppressors (preventing it), and their profiles differ meaningfully across cancer types.
Beyond miRNAs, long non-coding RNAs (lncRNAs) also play critical regulatory roles. Certain lncRNAs act as molecular sponges that compete with miRNAs for binding targets, creating complex regulatory networks known as competing endogenous RNA (ceRNA) networks that influence cancer progression, metastasis, and drug resistance.
Despite these known molecular relationships, most prior work had not fully exploited the intricate miRNA-mRNA-lncRNA interaction networks for pan-cancer classification. Earlier methods tended to look at each molecule type in isolation, missing the richer biological signal available from their combined interactions.
The research team obtained transcriptomic data from The Cancer Genome Atlas (TCGA), a large publicly available database of tumor samples. They focused on 14 cancer types, collecting both primary tumor tissue and matched normal tissue samples -- over 7,100 samples in total for miRNA sequencing data.
For each cancer type, they performed differential expression analysis using the DESeq2 statistical tool, comparing tumor versus normal tissue to identify miRNAs, messenger RNAs (mRNAs), and lncRNAs that were significantly up- or down-regulated. Only molecules with a fold-change of at least 2 and a statistically adjusted p-value below 0.05 were included.
The team then calculated Pearson correlation coefficients among the differentially expressed molecules to build co-expression networks for each cancer type. Only interactions with a correlation above 0.5 in absolute value were retained, resulting in a biologically meaningful miRNA-mRNA-lncRNA network per cancer. Across all 14 cancers, 597 unique interacting miRNAs were identified.
Network structure was further analyzed using community detection algorithms, and key properties like degree centrality -- a measure of how many connections a molecule has within the network -- were calculated. MiRNAs with the highest centrality, such as miR-145, miR-100-5p, and miR-143-3p, were found to be influential across multiple cancer types, suggesting broad roles in cancer biology.
Four ensemble machine learning (ML) algorithms were trained on the miRNA expression data: Random Forest (RF), AdaBoost, XGBoost, and LightGBM. Each algorithm builds predictions by combining many smaller decision models, which improves accuracy and reduces the risk of overfitting to the training data.
To handle the challenge of class imbalance (some cancer types had fewer samples than others), the team applied SMOTE (Synthetic Minority Over-sampling Technique), which generates synthetic data points for underrepresented classes. Data were split into 70% for training and 30% for testing, and a stratified 5-fold cross-validation approach ensured robust performance evaluation.
Four feature selection methods were used to narrow down the most informative miRNAs from the full set of 597: Recursive Feature Elimination (RFE), which iteratively removes the least useful features; the Boruta method, which identifies features better than random chance; Linear Discriminant Analysis (LDA); and Random Forest feature importance scoring.
Each method produced a different subset of miRNAs: RFE selected 150, RF selected 298, LDA selected 352, and Boruta retained 530. These subsets were then each used to train and test all four ML models, creating 25 total model-feature combinations to compare. Performance was assessed by accuracy, precision, recall, F1-score, and AUC.
All trained models performed remarkably well, achieving an overall 99% classification accuracy across 14 cancer types and 27 classes (tumor and normal tissue counted separately per cancer). This level of accuracy is among the highest reported for multi-cancer tissue-of-origin classification using transcriptomic data.
The RFE feature set with just 150 miRNAs proved optimal, achieving 99% accuracy with both macro and weighted averages at 99%. This is notable because it uses far fewer features than most competing methods while maintaining equivalent or superior performance, which has implications for clinical practicality and cost.
The ensemble voting classifier, combining all four ML algorithms, achieved an average accuracy of 99.03%. Notably, all five model types -- RF, AdaBoost, XGBoost, LightGBM, and the ensemble -- performed consistently, suggesting that the feature quality (driven by the network-based miRNA selection) was the dominant factor in success rather than the specific algorithm used.
A few cancer classes proved more challenging: bladder urothelial carcinoma normal tissue (BLCA-NT) had lower recall in some models (as low as 67%), and lung adenocarcinoma normal tissue (LUAD-NT) also saw reduced recall in certain configurations. This likely reflects molecular overlap between these normal tissue types and similar cancer types.
Feature importance analysis across all four ML models identified a consistent set of influential miRNAs. These included miR-21-5p, miR-93-5p, miR-10b-5p, miR-204-5p, miR-105-5p, and miR-139-5p -- molecules implicated in processes like immune modulation, epithelial-mesenchymal transition, angiogenesis, and chemotherapy resistance.
Survival analysis using Kaplan-Meier plots revealed strong prognostic associations. For example, high expression of miR-204-5p in breast cancer correlated with improved survival, while low expression of miR-105-5p in the same cancer was linked to poorer outcomes. In kidney renal clear cell carcinoma, elevated miR-10b-5p and miR-139-5p predicted better survival.
The identified miRNAs were cross-referenced against literature databases, cancer miRNA census (CMC) data, and clinical trial registries. A total of 159 predictive miRNAs matched published cancer literature entries, 126 were found in extracellular vesicle (EV) miRNA databases -- relevant for blood-based liquid biopsy detection -- and 202 matched the CMC list.
Several miRNAs are actively being investigated in ongoing clinical trials, including miR-10b-5p (glioblastoma), miR-155-3p/5p (lymphoma and breast cancer), miR-34a-5p (renal, lung, and liver cancer), and miR-193a-3p (advanced solid tumors). This overlap between computational predictions and clinical investigation validates the translational relevance of the study's findings.
Gene Ontology (GO) and KEGG pathway enrichment analyses were conducted on the experimentally validated targets of the 597 interacting miRNAs. This approach connects the molecular findings back to known biological processes and disease pathways, helping to explain why these miRNAs are so diagnostically informative.
The most significantly enriched KEGG pathways included cellular senescence, Hippo signaling, FoxO signaling, MAPK signaling, and TNF signaling -- all of which are established hallmarks of cancer biology. Enrichment of HPV-related pathways was also found, consistent with the known link between HPV infection and several gynecological and head-and-neck cancers.
GO biological process terms pointed to cancer-relevant functions including T-cell differentiation, DNA replication, cellular adhesion, and embryonic organ development. These terms reflect the dual nature of these miRNAs: regulating both normal tissue maintenance and the aberrant processes that drive cancer progression.
When comparing miRNA-based classifiers to those built using mRNA or lncRNA features, the miRNA models consistently outperformed -- especially for classifying normal tissue samples in cancer types like bladder, stomach, and prostate. This confirms that miRNA expression patterns carry particularly specific and robust tissue-of-origin information.
The study's framework has direct translational implications for cancer diagnostics. By identifying a compact set of 150 miRNAs capable of distinguishing 14 cancer types with 99% accuracy, the authors have laid groundwork for a potential multi-cancer detection test that could be applied in clinical settings.
Many of the identified miRNAs are detectable in blood through extracellular vesicles (EVs), making them strong candidates for non-invasive liquid biopsy approaches. Liquid biopsy -- detecting cancer-related molecules in blood rather than from tumor tissue -- could enable earlier detection at a stage when treatment is most effective.
The identification of miRNAs linked to drug sensitivity and resistance adds another clinical dimension. Among the 617 miRNAs assessed for drug associations, 186 were linked to validated treatment resistance and 157 to treatment sensitivity. This information could help guide personalized therapy decisions by flagging tumors likely to resist or respond to specific drugs.
The authors acknowledge that the study relied exclusively on TCGA data, which may not fully capture diversity across all patient populations or cancer stages. Future work will need to validate findings in independent cohorts, develop experimental confirmations of the network interactions, and explore whether the approach can be extended to rarer cancer types or cancers of unknown primary origin -- a particularly pressing clinical need.