Locally advanced rectal cancer (LARC) affects many patients worldwide, and the standard treatment involves a combination of chemotherapy and radiation given before surgery, known as neoadjuvant chemoradiotherapy (nCRT). This approach helps shrink tumors before removal, but the degree to which patients respond varies greatly from person to person.
Only a minority of patients achieve what doctors call a pathologic complete response, meaning the tumor disappears entirely after treatment. Patients who do not respond well not only miss out on benefits but also suffer unnecessary side effects from treatments that are not helping them.
Traditionally, doctors can only assess how well treatment worked after it is already complete, using MRI or CT scans and then examining the surgically removed tissue. This delayed evaluation makes it impossible to change course mid-treatment for patients who are not responding. There is a clear need for tools that can predict treatment response before therapy begins.
This study aimed to address that gap by using advanced CT imaging analysis combined with blood-based tumor markers and machine learning to predict which patients with LARC will respond well to nCRT before treatment even starts.
Radiomics is a technique that extracts a very large number of detailed measurements from standard medical images, such as CT scans, going far beyond what a radiologist can see with the naked eye. These measurements capture information about a tumor's shape, texture, and internal complexity, all of which reflect underlying biological characteristics.
When combined with machine learning (computer programs that learn patterns from data), radiomics has shown promise for predicting how cancers will respond to treatment. However, many machine learning models are considered "black boxes" because it is hard to understand why they make their predictions, which makes doctors hesitant to trust them.
To address this, researchers used a technique called SHAP (SHapley Additive exPlanations), which explains the contribution of each measurement to a model's prediction. This transparency is critical for doctors to feel confident in using an AI-assisted tool in real clinical settings.
The study also incorporated standard clinical biomarkers, specifically blood levels of CEA (carcinoembryonic antigen) and CA19-9, two proteins that are often elevated in colorectal cancer and can indicate tumor aggressiveness.
This was a retrospective study, meaning researchers looked back at medical records of patients already treated. A total of 272 patients with confirmed locally advanced rectal cancer were included from two hospitals in Chongqing, China, treated between 2021 and 2023.
Patients were divided into three groups: a training set of 156 patients (used to build the models), an internal validation set of 67 patients from the same hospital, and an external validation set of 49 patients from a different hospital. Using an external set from a different institution helps show whether the model works broadly, not just at one site.
All patients received a standardized treatment: radiation of 50.4 Gy delivered in 28 fractions over about six weeks, plus concurrent oral chemotherapy with capecitabine. After treatment, surgery was performed and tumor tissue was examined under a microscope using a grading system (TRG 0-3) to classify patients as responders (grade 0-1, good response) or non-responders (grade 2-3, poor response).
Pre-treatment CT scans were analyzed to extract 107 radiomic features covering shape, first-order statistics, and texture. Blood levels of CEA and CA19-9 were also recorded. Five different machine learning algorithms were then tested to find the best-performing model.
From 107 extracted features, researchers first checked how consistently each feature was measured between two different radiologists. Features with high consistency (ICC greater than 0.75) were kept, leaving 88 reliable features. A statistical technique called LASSO regression then narrowed this down to just 10 key features that most strongly predicted treatment response.
These 10 features were combined into a single number called the R-score (Radiomics Score), ranging from 0 to 1. A higher R-score indicated a greater likelihood of responding well to treatment. The R-score was significantly higher in responders than non-responders across all three patient groups.
Clinical variables like CEA and CA19-9 were also analyzed separately, and those significantly associated with response were included in a clinical model. The R-score was then combined with CEA and CA19-9 to create a combined model, which was tested using five different algorithms: logistic regression, random forest, support vector machine, decision tree, and XGBoost.
The XGBoost combined model outperformed all others in the training set and was selected as the final model for evaluation. Model accuracy was measured by the AUC (area under the ROC curve), where values closer to 1.0 indicate better predictive power.
The combined XGBoost model achieved an AUC of 0.844 in the internal validation set and 0.800 in the external validation set. These numbers show the model can distinguish responders from non-responders with strong accuracy. By comparison, the clinical model (using only CEA and CA19-9) achieved AUCs of only 0.747 to 0.768, and the imaging-only model reached 0.781 to 0.796.
The difference in performance between the combined model and single-component models was statistically significant (p less than 0.001 for all comparisons), meaning the improvement was not due to chance. This confirms that combining CT imaging data with blood biomarkers provides a richer and more accurate picture of likely treatment response than either approach alone.
The SHAP analysis confirmed that the R-score contributed the most to the model's predictions, far outweighing the contributions of CEA and CA19-9. Higher R-score values consistently predicted a greater probability of responding well to treatment. This result is both biologically meaningful and practically useful for clinicians.
Decision curve analysis, which evaluates whether a model provides net clinical benefit over a range of decision thresholds, showed the combined model consistently outperformed both the clinical and imaging models. This means using the combined model in practice would lead to better patient care decisions across a wide range of clinical scenarios.
For patients with locally advanced rectal cancer, this kind of predictive tool could transform how treatment decisions are made. Patients predicted to respond well could be offered organ-preservation strategies, potentially avoiding the need for extensive surgery if the tumor disappears completely after treatment.
Patients predicted not to respond well could be identified early and offered alternative or intensified treatments, rather than going through weeks of treatment that may not help them. This shift toward personalized treatment is often called precision medicine.
The model's consistency in the external validation set (from a different hospital) suggests it may generalize across institutions, an important quality for a tool that might eventually be deployed widely. The model does not require specialized equipment beyond standard CT imaging and routine blood tests, making it potentially accessible.
The inclusion of SHAP explainability means doctors can understand why the model predicts what it does, building the trust needed to actually use AI tools in clinical practice. The researchers also noted that a web-based or hospital system-integrated version of the tool would further support real-world adoption.
The authors acknowledge several important limitations. Because the study was retrospective, it relied on past records rather than prospectively planned data collection, which can introduce selection bias. The external validation set included fewer than 50 patients, which limits how confidently the findings can be generalized.
The study only used CT imaging, not MRI, which is also commonly used for rectal cancer staging and may capture additional information. Future studies should incorporate MRI features to compare approaches and explore whether combining modalities improves performance further.
Tumor outlines were drawn manually by radiologists, which introduces variability between observers. Automated segmentation tools could reduce this variability in future work. The study also did not include newer biomarkers such as circulating tumor DNA or immune profiling, which may add further predictive value.
Perhaps most importantly, the study measured short-term treatment response, not long-term outcomes such as whether patients lived longer or remained cancer-free. Future prospective studies with longer follow-up are needed to confirm whether predicting treatment response with this model also predicts improved survival outcomes.
This study successfully developed and validated a machine learning model that combines CT radiomics features with clinical biomarkers CEA and CA19-9 to predict how locally advanced rectal cancer patients will respond to pre-surgery chemoradiotherapy. The XGBoost-based combined model achieved strong predictive performance across multiple validation groups from different hospitals.
The use of SHAP explainability addresses one of the biggest barriers to AI adoption in medicine: the lack of transparency. By showing which features matter most and in what direction, the model can support rather than replace clinical judgment.
For patients and families, this research represents progress toward a future where cancer treatment decisions can be tailored to individual biology rather than applied uniformly. Though further validation in larger prospective studies is required before clinical adoption, this work provides a strong and reproducible foundation.
Ultimately, tools like this could help ensure that patients receive the right treatment from the start, sparing those with good responses from unnecessary surgery, and ensuring those less likely to respond are offered more aggressive alternatives earlier in their care.