Colorectal cancer is the third most commonly diagnosed cancer worldwide, responsible for nearly 916,000 deaths annually. When the cancer spreads to other organs - a stage called metastatic colorectal cancer (mCRC) - the five-year survival rate drops below 15%, making effective treatment selection critically important.
For patients with mCRC, chemotherapy regimens based on oxaliplatin are among the most commonly prescribed. These include combinations known as FOLFOX (oxaliplatin with leucovorin and fluorouracil) and XELOX (oxaliplatin with capecitabine). However, oxaliplatin can cause significant peripheral neuropathy - nerve damage that affects sensation and movement - making it important to identify which patients will actually benefit before prescribing it.
A major obstacle in clinical practice is the lack of reliable predictive biomarkers - measurable biological signals that can tell doctors in advance how a patient will respond to treatment. Existing approaches based on mRNA gene expression require fresh tumor tissue and complex handling, while DNA-based markers using inherited genetic variants have not been reliably validated in clinical settings.
Cancer cells frequently acquire large-scale changes in their DNA copy number - meaning segments of the genome are duplicated or deleted more times than normal. These changes, called copy number alterations (CNAs), are among the most prevalent genomic abnormalities in human cancers.
Unlike small mutations affecting a single DNA letter, CNAs involve stretching or shrinking large chromosomal regions. The pattern of these alterations across the entire genome can function as a kind of genomic fingerprint that reflects the history of a tumor's development and its underlying biological behavior.
Despite being widespread in cancer, CNAs have been underutilized as clinical biomarkers. This is partly because extracting meaningful patterns from the noisy, complex CNA landscape requires sophisticated computational approaches. This study addresses that gap by developing a machine learning model that converts a tumor's CNA profile into a predictive signal for chemotherapy response.
The researchers collected tumor tissue samples from 297 metastatic colorectal cancer patients treated at two major Chinese hospitals: Fudan University Shanghai Cancer Center (FUSCC) and Tongji Hospital. All samples were obtained before chemotherapy began, and only patients with documented oxaliplatin-based treatment and recorded clinical response data were included.
Rather than using expensive deep sequencing, the team chose shallow whole-genome sequencing (sWGS) - a cost-effective approach that provides enough genomic information to reliably detect copy number changes across the entire genome. This matters for clinical translation because lower cost makes broader patient access more feasible.
The 297 patients were divided into a training cohort used to build the model and three independent test cohorts used to validate it. This multi-cohort design is considered rigorous because it tests whether the model generalizes beyond the data it was trained on - a common weakness in many machine learning studies.
From each patient's sequencing data, the researchers computed 310 CNA features - quantitative measurements describing different aspects of genomic copy number patterns, such as the proportion of the genome affected by copy number changes, the distribution of segment sizes, and the intensity of amplifications on specific chromosomes.
An internal benchmarking process compared multiple machine learning algorithms and CNA feature subsets to find the best-performing combination. The winning approach was an XGBoost model - a powerful gradient-boosting algorithm - trained on just 7 key CNA features. XGBoost is known for its ability to handle complex, nonlinear data patterns while remaining resistant to overfitting.
This compact 7-feature model was named the CNA fingerprint. A classification threshold of 0.58 was selected to separate predicted responders from predicted non-responders. The model outputs a single score that can be interpreted directly by clinicians to inform treatment decisions.
The CNA fingerprint was evaluated on all three independent test cohorts. It achieved an area under the ROC curve (AUC) of 0.87 in both the first and second test cohorts, and 0.85 in the third cohort. An AUC near 1.0 indicates a near-perfect classifier, so values above 0.85 are considered strong performance for a biological prediction problem.
Crucially, the model's performance was consistent across cohorts from two different hospitals where samples were processed and sequenced independently. This robustness is meaningful because it suggests the CNA fingerprint is capturing a genuine biological signal rather than an artifact of a specific dataset or laboratory pipeline.
As a further validation, the researchers applied the CNA fingerprint to an external public dataset (GSE36864) containing 349 untreated mCRC patients receiving various chemotherapy regimens. Patients predicted as responders who actually received oxaliplatin-based therapy had better progression-free survival, while the same predictive advantage was not seen for other chemotherapy types - confirming specificity to oxaliplatin.
Among the seven features in the CNA fingerprint model, one feature stood out as the most important contributor to predictions: CN[>8]. This feature counts the number of DNA segments where the absolute copy number exceeds 8 - meaning those segments have been duplicated to an extreme degree compared to the normal two copies per cell.
Patients who responded well to oxaliplatin-based chemotherapy consistently had lower CN[>8] counts compared to non-responders across both the training cohort and all three test cohorts. In other words, tumors with less extreme genomic amplification tended to be more sensitive to oxaliplatin treatment.
The biological meaning of this finding is still being investigated, but one hypothesis involves genomic instability and DNA repair. Highly amplified genomes may have acquired increased capacity to repair DNA damage - which is essentially how oxaliplatin kills cells. Tumors with this repair advantage could be more resistant to the drug's effects, explaining the correlation between high CN[>8] and treatment failure.
Kaplan-Meier survival analysis showed that patients predicted as responders by the CNA fingerprint achieved significantly better overall survival compared to predicted non-responders after receiving oxaliplatin-based chemotherapy. This survival benefit was not observed in patients who received different chemotherapy regimens, adding further evidence of specificity.
The researchers compared the CNA fingerprint against several established or proposed biomarkers: KRAS mutation status, aneuploidy score, primary tumor location (left vs. right colon), and HRD (homologous recombination deficiency) status. None of these alternatives could reliably predict oxaliplatin response in this patient population.
Notably, HRD status - which is a validated biomarker for platinum-based chemotherapy in ovarian cancer - showed no predictive value for oxaliplatin response in colorectal cancer. This highlights that biomarkers do not always transfer between cancer types, and underscores the need for CRC-specific tools like the CNA fingerprint.
One practical advantage of the CNA fingerprint is that the underlying data - a genome-wide copy number profile - can be generated using shallow whole-genome sequencing, which is considerably more affordable than standard deep sequencing approaches. This makes the test economically feasible for real-world clinical deployment.
To lower the barrier to adoption, the researchers created an open-source R software package that takes a standard CNA profile file as input and outputs the CNA fingerprint score for each patient. Clinicians and researchers can apply the tool without requiring specialized bioinformatics expertise.
The study acknowledges important limitations: all data were collected retrospectively, cohort sizes were relatively small, and the distribution of patient characteristics was not always balanced. The authors are initiating prospective clinical trials to formally evaluate whether the CNA fingerprint can guide real-time treatment decisions in newly diagnosed mCRC patients, which would be the necessary step toward regulatory approval and clinical adoption.
This study demonstrates for the first time that CNA-based features can serve as actionable biomarkers for predicting chemotherapy drug response - a concept that had been largely unexplored compared to mutation- or expression-based approaches. The CNA fingerprint approach opens a new class of genomic biomarkers for precision oncology.
Because the sWGS method used to generate CNA profiles is applicable to virtually any tumor type, the same general strategy could be adapted for other cancers and other chemotherapy drugs. The authors suggest that similar CNA fingerprint biomarkers could be developed for breast cancer, gastric cancer, and other malignancies where platinum-based agents are used.
For patients with metastatic colorectal cancer, this research represents a step toward a future where a simple, affordable genomic test on pre-treatment tumor tissue can reliably identify who should receive oxaliplatin-based therapy - sparing non-responders from side effects like neuropathy while concentrating treatment where it is most likely to extend life.