Prostate cancer is the third most prevalent cancer in men worldwide, accounting for roughly 7% of all cancer incidence and nearly 4% of cancer-related deaths globally. Despite its prevalence, diagnosing and managing it accurately remains a significant challenge.
The current gold-standard screening tool, prostate-specific antigen (PSA), is an organ-specific rather than tumor-specific marker. This means elevated PSA can be caused by non-cancerous conditions, leading to unnecessary biopsies and overdiagnosis of low-risk tumors that may never harm the patient.
Biopsy is an invasive procedure with real risks, and existing supplementary tests like PHI, PCA3, and 4Kscore have not definitively surpassed PSA in clinical utility. There is a pressing need for novel, blood-based protein biomarkers that are more specific to prostate cancer.
Proteomics -- the large-scale study of proteins -- offers a promising avenue. Proteins are direct functional players in disease progression and serve as active therapeutic targets, making them ideal candidates for biomarker discovery beyond what genomics or transcriptomics can offer alone.
SWATH-MS (Sequential Window Acquisition of All Theoretical Fragment-Ion Spectra mass spectrometry) is a cutting-edge data-independent acquisition technique capable of detecting and measuring thousands of proteins in a biological sample simultaneously.
Unlike traditional targeted methods, SWATH-MS creates permanent digital records of all peptide fragments within a defined mass range. These records can be re-analyzed at any time as long as a spectral reference library exists to match the fragments against.
The quality and comprehensiveness of this reference library is critical -- it determines how many proteins can be reliably identified and quantified in future experiments. Without a disease-specific library, many relevant proteins in cancer samples go undetected.
Prior prostate cancer spectral libraries focused on tumor tissue or cancer cell lines, not blood. Since blood-based, non-invasive diagnosis is the clinical goal, a dedicated serum spectral library for prostate cancer was a significant unmet need.
The study enrolled 373 prostate cancer patients spanning low, intermediate, and high-grade disease, as well as 134 healthy male controls with normal PSA levels. Blood serum was collected, processed, and pooled to create the discovery dataset.
To uncover low-abundance proteins hidden beneath highly abundant blood proteins like albumin, the top 12 most abundant proteins were removed using immunodepletion columns. Samples were then subjected to deep fractionation to maximize the number of detectable proteins.
Samples were analyzed using three different mass spectrometry platforms: an Orbitrap Fusion Lumos and a Sciex TripleTOF 6600 in both micro-flow and nano-flow modes. This multi-instrument approach ensured broader coverage of the proteome than any single instrument could provide.
A fourth data source, an in silico disease library, was created by pulling known prostate cancer proteins from the Jensen Disease Database and generating theoretical peptide spectra. All four libraries were then merged into one comprehensive combined library using a web-based tool called iSwathX.
The final combined prostate cancer serum spectral library contains 114,684 transitions, 18,479 peptides, and 1,227 unique proteins. This represents the most comprehensive blood-based proteomics resource for prostate cancer created to date.
Each individual library contributed unique proteins. The Orbitrap Lumos library identified 760 proteins, while the Sciex nano-flow and micro-flow libraries identified 712 and 484 proteins respectively. The nano-flow approach proved more thorough than micro-flow, though micro-flow is faster and better suited for clinical throughput.
The proteins in the library are predominantly extracellular proteins involved in regulatory, structural, and binding functions -- consistent with proteins expected in blood serum. Gene ontology analysis confirmed their roles in cellular metabolism, developmental processes, and signaling pathways.
Compared to the Pan-Human spectral library covering over 10,000 human proteins, the combined prostate cancer library contained 154 proteins not found in the pan-human resource. These unique proteins were enriched in cancer and hormonal signaling pathways, and included known prostate-specific markers like KLK2, KLK11, TMEFF2, and ANO7.
Among the 154 proteins unique to the prostate cancer serum library, several are encoded by well-known cancer-associated genes including TNF, MYC, BRCA2, and ERG1. These proteins are implicated in tumor growth, DNA repair, and aggressive disease behavior.
Cross-referencing the library with the COSMIC database (Catalogue of Somatic Mutations in Cancer) revealed 64 overlapping genes, demonstrating that many proteins in the blood library are directly connected to cancer-driving mutations.
KLK2 and KLK11, both kallikrein-related peptidases, were identified as prostate-specific proteins in the library. These are structurally related to PSA (KLK3) and represent potential candidates to supplement or replace PSA-based testing.
Testis-specific proteins such as BRCA2, TERT, and ESR2 were also present, reflecting the reproductive tissue origins of many prostate cancer-relevant proteins. The diversity of tissue associations underscores the complexity of the prostate cancer blood proteome.
To test whether the library could be used reliably on new patients, an independent cohort of 29 pre-treatment and 14 post-radiotherapy prostate cancer patients was analyzed. Their blood samples were not used during library construction, ensuring an unbiased validation.
Using SWATH-MS guided by the new library, 404 proteins were reliably identified at a strict 1% false discovery rate. The library successfully guided protein identification across a dynamic range spanning six orders of magnitude -- far exceeding the 2-4 log range typical of serum proteomics studies.
A random forest machine learning model applied to the protein data cleanly separated pre-treatment from post-treatment patient samples. The 17 most informative proteins included APOM, TTR, APOH, and SAA4, several of which are involved in lipid transport and acute-phase immune responses.
When highly variable or non-significant proteins were removed, the remaining 12-protein signature still produced clear clustering of treatment groups on a heatmap analysis, confirming that the library can detect biologically meaningful protein changes associated with cancer treatment.
The study tested whether six previously reported prostate cancer biomarker proteins could be detected and measured using the new library. All six -- KLKB1, MYC, IGF1, SRC, BRCA1, and STAT3 -- were successfully quantified across validation patient samples.
KLKB1 (plasma kallikrein, a PSA-related serine protease) showed the highest average expression intensity, consistent with its known role in prostate cancer. MYC overexpression correlates with disease severity, while IGF1 promotes tumor formation via downstream growth signaling.
STAT3 and its activator, the IL-6 receptor, are linked to prostate cancer progression and represent potential therapeutic targets. The library's ability to measure all these proteins simultaneously demonstrates its breadth as a biomarker discovery tool.
The library also detected the three proteins that bind to free PSA in the bloodstream -- alpha-1 antitrypsin, alpha-1 antichymotrypsin, and alpha-2 macroglobulin -- with mean expression levels quantified across all validation samples, expanding the utility of PSA-related measurements.
Human serum is dominated by a small number of highly abundant proteins like albumin and globulins, which account for 99% of total protein content. These mask the low-abundance proteins most likely to serve as disease markers, requiring careful depletion strategies before meaningful proteomics analysis is possible.
Previous SWATH reference libraries for prostate cancer were built from tumor tissue or cell lines, which are poor surrogates for blood. Tissue biopsies are invasive, and cell lines may not accurately reflect the biology of actual tumors. Blood-based libraries directly address the clinical goal of non-invasive testing.
By comparing the combined prostate cancer serum library to two existing tissue/cell-line libraries, the study found 959 proteins unique to the serum library -- confirming that blood proteomics captures a fundamentally different and complementary picture of the disease.
The shift in clinical research from single biomarkers to multi-protein panels is well-supported by this library approach. Rather than relying on PSA alone, future tests may draw on panels of blood proteins identified and quantified using SWATH-MS guided by resources like this one.
The combined prostate cancer serum spectral library -- with 1,227 proteins and nearly 115,000 spectral transitions -- is the most comprehensive blood-based proteomics resource for prostate cancer available. It is publicly accessible through the ProteomeXchange database (identifier PXD028651) for use by other researchers.
Unlike tissue or cell-line libraries, this resource spans the full clinical spectrum from low-grade indolent to high-grade metastatic prostate cancer, making it applicable to research across disease stages and treatment contexts.
The library can be combined with existing tissue reference libraries to provide researchers with a complete view of prostate cancer proteome dynamics across different biological compartments -- from the tumor itself to what appears in the bloodstream.
While SWATH-MS itself may not become a routine clinical test overnight, the protein panels it identifies can pave the way for simpler, more affordable immunoassay-based tests that could one day replace or meaningfully supplement PSA for prostate cancer diagnosis and monitoring.