Prostate-specific antigen (PSA) testing has been the cornerstone of prostate cancer screening for decades, but it suffers from poor specificity. High PSA can reflect benign enlargement, inflammation, or other non-cancerous conditions, leading to unnecessary biopsies and overtreatment.
Current clinical staging tools such as the Gleason score and PSA serum levels guide treatment decisions, but they cannot reliably distinguish indolent from aggressive tumors at the molecular level, a gap that precision medicine aims to close.
A growing number of commercially available molecular tests, including OncotypeDX Genomic Prostate Score, Prolaris, ProMark, and Decipher, predict recurrence risk based on cancer-associated gene panels. However, each works for only a subset of patients and none covers the full spectrum of prostate cancer biology.
This systematic review synthesizes the latest research on coding genes, non-coding RNAs, repetitive sequences, sequencing technologies, and artificial intelligence as a roadmap toward better, more personalized biomarkers for prostate cancer management.
Several protein-coding genes have emerged as clinically relevant biomarkers. TMPRSS2-ERG, a gene fusion present in roughly 50 percent of prostate tumors, is being evaluated in multiple Phase I and II clinical trials as a diagnostic and therapeutic target. The androgen receptor (AR) and its splice variant AR-V7 are strongly linked to resistance against enzalutamide and abiraterone in castration-resistant disease.
BRCA2 mutations, found in approximately 12 percent of castration-resistant prostate cancer cases, predict sensitivity to PARP inhibitor therapy such as olaparib. Likewise, PTEN loss activates the PI3K/AKT/mTOR pathway and is associated with aggressive, therapy-resistant disease.
Epigenetic regulators including MGMT, DNMT1, and JMJD3 affect DNA methylation and have been linked to prostate cancer mortality and tumor development, while splicing factors such as CDK9 and SF3B2 influence the generation of the AR-V7 variant and are emerging targets for therapeutic intervention.
Multi-gene germline panels identify pathogenic variants in 7 to 12 percent of prostate cancer patients across DNA repair genes including BRCA1, BRCA2, ATM, and mismatch repair genes MLH1, MSH2, and MSH6, findings that have direct implications for treatment selection and family cancer risk counseling.
MicroRNAs (miRNAs) are short RNA molecules of 21 to 25 nucleotides that regulate gene expression post-transcriptionally. Because they are stable in body fluids such as blood, urine, and semen, and show tissue- and stage-specific expression patterns, they are attractive liquid biopsy biomarkers.
Specific miRNAs including miR-21, miR-221, miR-1290, and miR-375 are overexpressed in castration-resistant prostate cancer and correlate with prognosis. A four-miRNA panel (miR-4289, miR-326, miR-152-3p, and miR-98-5p) distinguished prostate cancer patients from healthy controls with an area under the ROC curve of 0.88.
Long non-coding RNAs (lncRNAs) are transcripts longer than 200 nucleotides with no protein-coding potential. They regulate chromatin remodeling, gene transcription, and mRNA stability. PCA3 was the first lncRNA to receive FDA approval (2012) as a prostate cancer biomarker through the PROGENSA test, reducing unnecessary biopsies by roughly 40 percent.
SChLAP1 lncRNA antagonizes the SWI/SNF chromatin remodeling complex, promoting metastasis, and is one of the strongest predictors of biochemical recurrence and metastasis in prostate cancer patients. PCAT1 drives tumor progression via the Wnt/beta-catenin pathway and negatively regulates the BRCA2 tumor suppressor, making it a dual diagnostic and prognostic candidate.
Repetitive sequences make up approximately 55 percent of the human genome. Historically overlooked, advances in whole-genome and transcriptome sequencing have made it possible to detect and quantify these elements systematically across cancer samples.
HERV-K, a human endogenous retroviral sequence, is highly expressed in malignant prostate tissue compared with normal tissue. It is detectable in blood and may enhance the diagnostic efficiency of PSA testing when used in combination.
LINE-1, a transposable element encoding the RNA-binding protein ORF1p, shows increased expression in prostate cancer tissue. Its hypomethylation is associated with tumor progression, providing a potential epigenetic biomarker of disease stage.
Although the evidence base for repetitive sequence biomarkers is smaller than for miRNAs and lncRNAs, these elements represent an underexplored layer of molecular information that could improve diagnostic sensitivity when combined with other biomarker types.
RNA-Seq (transcriptome sequencing) captures the expression levels of all coding and non-coding genes simultaneously, enabling the discovery of novel transcripts, gene fusions, differential splicing, and expression changes that classic methods such as qPCR cannot detect in an unbiased, genome-wide manner.
Spatial transcriptomics maps gene expression to specific tissue locations within a prostate biopsy, revealing the molecular heterogeneity of tumors. Studies using this technique identified cancer-specific markers such as SPINK1 and PGC, stroma-specific markers like NR4A1, and PIN-specific markers such as NPY, demonstrating the value of spatial resolution in biomarker discovery.
Exome and whole-genome sequencing have uncovered 70 previously unknown significantly mutated genes in prostate cancer, including CUL3, SPEN, and the epigenetic regulators KMT2C and KMT2D. Whole-genome sequencing also identified complex chromosomal rearrangements called chromoplexy, a pattern of coordinated multi-gene disruptions that support a punctuated model of cancer evolution.
Exome sequencing of castration-resistant prostate cancer patients showed actionable mutations in 90 percent of samples, with BRCA2 mutations in 12 percent and ATM mutations in 22 percent of cases, directly guiding PARP inhibitor treatment decisions and demonstrating the clinical value of comprehensive genomic profiling.
Machine learning (ML) algorithms, including supervised approaches such as logistic regression and random forest and unsupervised approaches such as principal component analysis, can detect patterns in large genomic and imaging datasets that are impossible to identify manually.
The commercial genomic classifier Decipher uses a random forest algorithm trained on the expression of 22 RNA biomarkers to predict metastasis risk after radical prostatectomy. A support vector machine model trained on 29 miRNAs achieved 97 percent accuracy in diagnosing prostate cancer, while a 7-miRNA prognostic model reached 66 percent accuracy for outcome prediction.
Deep learning applied to histological slides has produced models that can recognize tumor tissue across 400 slides from different patients and generate three-dimensional reconstructions of prostate cancer architecture to improve Gleason grading beyond what pathologists can achieve by visual inspection alone.
AI algorithms are also being used to identify mutations that cause treatment resistance. A deep neural network trained on prostate cancer genomic data successfully predicted which AR mutations confer resistance to the drug darolutamide, opening the door to AI-guided therapy selection that adapts to the evolving molecular landscape of each patient's tumor.
The review establishes that the combination of coding genes, non-coding RNAs, repetitive sequences, and AI-based analysis represents the most promising path toward precision medicine for prostate cancer, replacing the one-size-fits-all PSA screening paradigm with a multi-layered molecular portrait of each patient's disease.
Key biomarker targets with the strongest clinical evidence include TMPRSS2-ERG, AR-V7, BRCA2, PCA3, SChLAP1, miR-21, and LINE-1, most of which are already in clinical trials or FDA-approved tests. Their combined use in multi-biomarker panels is expected to dramatically improve sensitivity and specificity over any single marker.
Emerging sequencing technologies, particularly spatial transcriptomics and single-cell RNA-Seq, will resolve the spatial and cellular heterogeneity of prostate tumors, revealing subpopulation-specific biomarkers that bulk tissue analysis misses and enabling more targeted biopsy and treatment strategies.
The authors conclude that machine learning will accelerate the discovery and validation of novel biomarkers and, combined with genomic profiling, will drive a scientific revolution in prostate cancer management, improving stratification between indolent and aggressive disease and ultimately improving patient quality of life.