Metabolomic machine learning-based model predicts efficacy of chemoimmunotherapy for advanced lung squamous cell carcinoma

Front Immunol 2025 AI 7 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Predicting Who Benefits from Chemoimmunotherapy in Lung Squamous Cell Carcinoma

The clinical challenge Lung squamous cell carcinoma accounts for about 30% of all non-small cell lung cancer (NSCLC) cases. Unlike adenocarcinoma, it rarely harbors actionable driver gene mutations, leaving patients with fewer targeted therapy options and making chemoimmunotherapy the standard first-line approach.

Why new biomarkers are needed Existing biomarkers such as PD-L1 expression and tumor mutational burden (TMB) have shown limited ability to predict which patients with squamous carcinoma will respond to combined immune checkpoint inhibitor and chemotherapy regimens. This creates a need for new, more reliable predictive tools.

Study approach Researchers collected serum samples from 79 patients with advanced lung squamous cell carcinoma before they started chemoimmunotherapy. Using untargeted metabolomics - measuring thousands of small molecules in blood - combined with machine learning, the team aimed to identify a metabolite panel that could predict treatment outcome and survival.

Key finding in brief A panel of 8 serum metabolites, analyzed by a random forest (RF) machine learning model, achieved an area under the curve (AUC) of 0.973 in the training set and 0.944 in the validation set - indicating strong ability to separate patients likely to respond from those who are not.

TL;DR: A blood-based metabolomics and machine learning approach identified 8 serum metabolites that can predict survival outcomes in patients with advanced lung squamous cell carcinoma receiving chemoimmunotherapy, outperforming traditional biomarkers.
Pages 2-3
Limitations of Current Predictive Biomarkers and the Promise of Metabolomics

PD-L1 and TMB limitations While PD-L1 and TMB are used clinically to guide immunotherapy decisions, studies such as CheckMate 017 and 078 showed that progression-free survival improved with chemoimmunotherapy regardless of PD-L1 level. TMB thresholds also remain non-standardized, limiting its utility as a biomarker.

Advantages of blood-based metabolomics Blood samples can be collected non-invasively and repeatedly, allowing dynamic monitoring of the tumor microenvironment. Unlike single-timepoint tissue biopsies, metabolites in peripheral blood can reflect real-time changes in tumor biology and patient immune status.

Machine learning and omics integration The combination of metabolomics with AI-based modeling has already been applied in cancer diagnostics - for example, building models to distinguish NSCLC from healthy controls with AUC values near 0.96. The authors sought to extend this approach to treatment response prediction specifically in squamous cell carcinoma.

Study novelty Previous metabolomics studies focused primarily on early cancer detection rather than immunotherapy response prediction. This study represents one of the first attempts to use untargeted serum metabolomics combined with multiple machine learning algorithms to predict chemoimmunotherapy outcomes in advanced squamous carcinoma.

TL;DR: PD-L1 and TMB are imperfect predictors of chemoimmunotherapy response, and blood-based untargeted metabolomics combined with machine learning offers a promising non-invasive alternative for identifying patients who will benefit.
Pages 3-5
Study Design: Patient Selection, Metabolomics Analysis, and Machine Learning Pipeline

Patient cohort A total of 79 patients with stage IIIB-IV lung squamous cell carcinoma, all driver gene-negative, were enrolled at Shanghai Chest Hospital. All received first-line PD-1 inhibitor plus platinum-based chemotherapy. Serum samples were collected within 4 weeks before treatment initiation.

Response grouping Patients were divided into Response (R, n=41, OS at least 24 months) and Non-Response (NR, n=38, OS less than 24 months) groups, reflecting a clinically meaningful threshold for durable benefit from chemoimmunotherapy.

Untargeted metabolomics workflow Serum samples underwent liquid chromatography-mass spectrometry (LC-MS) analysis on a TripleTOF 6600 instrument. Data quality was maintained with pooled quality control (QC) samples inserted every 10 injections, and features with more than 50% missing values were excluded. Metabolite identification used self-built libraries and the KEGG database.

Machine learning pipeline From 117 initially identified differential metabolites, LASSO regression narrowed the list to 46, and random forest (RF) feature selection identified an 8-metabolite panel. RF, support vector machine (SVM), and logistic regression (LR) models were trained on a 70:30 split, and ROC curves with 1000 bootstrap iterations were used for performance assessment.

TL;DR: The study used untargeted LC-MS metabolomics on pre-treatment serum from 79 patients, followed by a LASSO and random forest feature selection pipeline, to build and compare three machine learning models for predicting chemoimmunotherapy response.
Pages 5-8
Differential Metabolites and Model Performance

Metabolic differences between groups OPLS-DA analysis clearly separated NR and R patients based on serum metabolite profiles. A total of 117 differential metabolites were identified (VIP greater than 1, p less than 0.05), with 70 upregulated and 47 downregulated in the NR group. Key KEGG pathways enriched included chemical carcinogenesis-receptor activation and cortisol synthesis.

The 8-metabolite panel After LASSO and random forest intersection analysis, 8 metabolites were selected as the best biomarker panel: Mevalonate, 1,2-Ethanedithiol, 4alpha-Hydroxymethyl-4beta-methyl-5alpha-cholesta-8,24-dien-3beta-ol, Arg-Asp-Leu-Tyr-Ser, Donhexocin, Glu-Cys-Ala, TG(10:0/14:0/a-15:0), and D-3-Hydroxykynurenine. No single metabolite alone showed sufficient predictive power (AUC less than 0.75); their combination achieved AUC 0.915.

Model comparison Among RF, SVM, and LR models trained on the 8-metabolite panel, the RF model achieved the best overall performance: AUC 0.973 in training and 0.944 in validation, with accuracy, precision, recall, and F1 all at 0.98 in training and 0.96-1.0 in validation.

Survival stratification Using the median risk score as a cutoff, patients were divided into high-risk and low-risk groups. High-risk patients had significantly shorter median OS (18 months vs. longer in low-risk group, p=0.00023) and PFS (7 months vs. 15 months, p less than 0.0001). Multivariate Cox regression confirmed that risk score and age were independent prognostic factors.

TL;DR: The 8-metabolite random forest model distinguished responders from non-responders with AUC exceeding 0.94, and the derived risk score predicted survival independently from clinical factors, with high-risk patients experiencing significantly shorter OS and PFS.
Pages 9-12
What the Key Metabolites Reveal About Tumor Biology

Mevalonate pathway Mevalonate is a central intermediate in cholesterol biosynthesis and cellular metabolism. Research has shown that the mevalonate pathway regulates adaptive immunity and can serve as a vaccine adjuvant and immunotherapy drug target, suggesting its role in modulating antitumor immune responses in squamous carcinoma.

D-3-Hydroxykynurenine This tryptophan metabolism byproduct primarily functions in neurodegenerative disease contexts but is also implicated in apoptosis regulation. Its presence in the predictive model hints at immunosuppressive tryptophan catabolism pathways that may blunt antitumor immunity.

Donhexocin and other metabolites Donhexocin, like D-3-Hydroxykynurenine, has been linked to apoptosis. Several of the 8 metabolites are poorly characterized, highlighting the ability of untargeted metabolomics to uncover novel biology that targeted approaches would miss.

Panel-based approach rationale The eight metabolites showed minimal correlation with each other yet together achieved strong predictive performance. This reflects the biological complexity of the tumor microenvironment: no single biomarker captures the full picture, while a multi-marker panel can model the combined metabolic state predictive of treatment outcome.

TL;DR: Key predictive metabolites implicate the mevalonate pathway and tryptophan catabolism in determining chemoimmunotherapy outcomes, and the need for a multi-metabolite panel reflects the inherent biological complexity of the tumor immune microenvironment.
Pages 8, 9, 13
Clinical Value and Potential Applications of the Metabolite Model

Blood-based monitoring advantage The model uses pre-treatment peripheral blood serum, making it non-invasive, easy to collect, and capable of capturing whole-tumor metabolic states rather than the local snapshot provided by tissue biopsy. This could allow dynamic monitoring at multiple timepoints during treatment.

Independent prognostic value The risk score derived from the model served as an independent predictor of OS and PFS in multivariate Cox analysis, alongside age. Time-dependent ROC curves showed AUCs of 0.79, 0.90, and 0.73 for 1-, 2-, and 3-year OS prediction, supporting clinical utility across multiple time horizons.

Kit development potential The authors suggest that their 8-metabolite panel could be developed into a clinical testing kit for routine pretreatment evaluation of patients with advanced squamous carcinoma, enabling oncologists to identify those less likely to benefit from standard chemoimmunotherapy and consider alternative or intensified regimens.

Complementing PD-L1 testing Since PD-L1 and clinical factors failed to predict outcomes in this cohort, the metabolomics model offers a complementary - or potentially superior - biomarker strategy. Integration with existing clinical risk indices could further improve patient selection.

TL;DR: The metabolite risk score independently predicted survival outcomes and holds promise as a clinical tool to guide chemoimmunotherapy decisions in advanced lung squamous cell carcinoma, potentially as a companion diagnostic kit.
Page 13
Study Constraints and the Path to Clinical Translation

Retrospective single-center design The study analyzed data from 79 patients at a single institution (Shanghai Chest Hospital). This limits the generalizability of the findings and introduces potential selection biases. Larger, multi-center prospective studies are essential to confirm the model's real-world performance.

Lack of external validation While internal cross-validation was performed, the model was not tested on an independent external cohort from a different institution or population. External validation is critical to establish the model's robustness across different patient demographics, treatment protocols, and laboratory environments.

Mechanistic gaps Although 8 metabolites were identified as predictive biomarkers, the mechanistic pathways linking these specific molecules to immunotherapy resistance or sensitivity remain incompletely understood. Functional studies are needed to clarify whether these metabolites are causal drivers of treatment response or downstream markers.

Future directions Prospective multicenter trials incorporating this metabolomics panel alongside standard biomarkers (PD-L1, TMB) are needed. Integration with other omics layers (proteomics, genomics) and clinical variables could further refine prediction. Development of a standardized, clinically deployable assay kit represents the key translational milestone.

TL;DR: The study's single-center retrospective design and lack of external validation are important limitations; future prospective multicenter studies and mechanistic investigation are needed before this metabolomics model can be adopted in clinical practice.
Citation: Open Access, 2025. Available at: PMC12000773.