The landscape of conventional and artificial intelligence-based clinical prediction models in non-small-cell lung cancer: from development to real-world validation

ESMO Open 2025 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Clinical Prediction Models: What They Are and Why They Matter

Lung cancer remains the leading cause of cancer death worldwide. Non-small-cell lung cancer (NSCLC) is the most common subtype, and despite advances in detection and treatment, outcomes remain poor with 1.8 million deaths recorded in 2022. The complexity of NSCLC - spanning early resectable stages to advanced metastatic disease - requires nuanced tools to guide clinical decisions beyond broad staging categories.

Clinical prediction models help oncologists make evidence-based decisions. These models use combinations of patient data - clinical variables, biomarkers, imaging features, and genetic information - to estimate the likelihood of specific outcomes such as overall survival, disease recurrence, or treatment response. Rather than relying solely on clinical intuition, they provide quantified probability estimates that can personalize treatment recommendations.

Current limitations hold back wider adoption. Despite an increasing volume of published models, translation to routine clinical practice remains limited. Key barriers include insufficient external validation in independent patient populations, the retrospective nature of most datasets, and the tendency to develop models on specific institutional cohorts that may not generalize to other healthcare settings.

AI is transforming prediction model development. Machine learning, deep learning, and natural language processing are increasingly being applied to integrate diverse data types - CT and PET scans, radiomics features, histopathology images, genomic profiles, and clinical records - to build more accurate, comprehensive, and scalable prediction tools. This review maps the full landscape from conventional models to AI-enhanced approaches across both early-stage and metastatic NSCLC.

TL;DR: This comprehensive review maps the landscape of clinical prediction models in NSCLC, highlighting their potential to personalize treatment while identifying key barriers to real-world clinical adoption, particularly the lack of external validation.
Pages 2, 3, 5
Prediction Models in Early-Stage NSCLC

Early NSCLC is primarily treated with surgery, but risk stratification is underdeveloped. Current European Society for Medical Oncology guidelines center on complete surgical resection for stages I to IIIA, with adjuvant chemotherapy for stages IIB and III. However, risk stratification within these stages - to identify who needs additional treatment and who can safely avoid it - still relies predominantly on the TNM staging system, which is imprecise at the individual patient level.

Demographic and clinical models have identified key prognostic factors. A major prospective multicenter study of 2,052 early-stage NSCLC patients found that older age, male sex, weight loss greater than 5%, current smoking, and higher TNM stage significantly predicted worse overall survival. Stage IIIA patients had a five-fold higher mortality risk than stage I patients. Notably, tumor histology alone was not a consistent independent predictor in some studies.

Histopathological models add a layer of tissue-level prediction. The CellDiv model quantified nuclear morphological diversity in tumor epithelium using 11 features of nuclear shape and intensity. In independent validation, high-risk patients predicted by CellDiv had significantly worse five-year overall survival in both lung adenocarcinoma (hazard ratio 1.62) and squamous cell carcinoma (hazard ratio 2.24). These nuclear features correlated with anti-apoptotic signaling pathways, suggesting a biological basis for their prognostic value.

Gene expression adds prognostic power beyond clinical variables. An immune-related gene prognostic index (IRGPI) based on 25 gene pairs from TCGA and other datasets significantly stratified stage I and II NSCLC patients into high- and low-risk groups, with approximately 34 percentage point differences in five-year survival rates. Such gene-based models hold promise for identifying patients with early-stage disease who need closer surveillance or adjuvant therapy.

TL;DR: Early-stage NSCLC prediction models incorporate demographic, histopathological, and genomic features to improve risk stratification beyond the TNM staging system, though prospective validation remains an unmet need.
Pages 5-7
Prediction Models in Metastatic NSCLC

Blood-based indices are practical and accessible prognostic tools. Several models for metastatic NSCLC use routinely available blood test results to stratify patients by survival risk. The neutrophil-to-lymphocyte ratio (NLR) has emerged as a consistent prognostic marker - reflecting the balance between inflammation and immune surveillance - with elevated NLR predicting worse outcomes across multiple independent studies of patients on immunotherapy.

Three key immunotherapy-era prediction tools have been developed. The LIPS-3 score (NLR, ECOG performance status, and steroid use) predicted one-year overall survival ranging from 78.2% in favorable-risk to just 10.7% in poor-risk groups among first-line pembrolizumab patients. The LITE Risk score stratified patients into median overall survival of 28.3, 9.1, and 2.1 months across favorable, intermediate, and poor categories. The LIPI (derived NLR and lactate dehydrogenase) identified three groups with median OS of 34, 10, and 3 months for good, intermediate, and poor risk on checkpoint inhibitors.

Gene expression signatures stratify metastatic risk. Studies using transcriptomic data from public databases identified gene panels associated with metastatic potential. One six-gene signature classified patients into high- and low-risk groups for metastasis using differentially expressed genes enriched in epidermal cell development pathways. A separate 11-gene model specifically targeting brain metastasis prediction identified an immunophenotypic difference between risk groups, with high-risk patients showing elevated tumor-associated macrophages and monocytes - an immunosuppressive pattern that may paradoxically facilitate metastasis.

Radiomic CT features can non-invasively predict immunotherapy response. The Lung Cancer Immunotherapy-Radiomics Prediction Vector (LCI-RPV) integrated 15 radiomic features from CT imaging to predict PD-L1 expression (AUC 0.70), treatment response at 3 months (AUC 0.68), and pneumonitis risk (AUC 0.64). The model stratified patients into high- and low-risk groups for three-year overall survival with hazard ratios of 2.26 to 2.45 across two independent validation cohorts, offering a non-invasive alternative to biopsy-based biomarker testing.

TL;DR: Metastatic NSCLC prediction models range from simple blood-based indices like LIPS-3 and LIPI to sophisticated gene expression signatures and radiomic models predicting immunotherapy response, all demonstrating significant survival stratification but lacking prospective randomized validation.
Pages 7-8
AI Technologies Applied to NSCLC Prediction

Three main AI paradigms are used in NSCLC prognostics. Machine learning (ML) algorithms analyze structured patient data to identify patterns and risk factors. Deep learning (DL), a more powerful subset, uses multi-layered neural networks capable of learning complex features from raw images and genomic data. Natural language processing (NLP) extends AI to unstructured clinical text, enabling extraction of prognostic information from radiology reports, pathology notes, and electronic health records.

Deep learning consistently outperforms traditional staging. A DL survival neural network (DeepSurv) trained on 17,322 NSCLC patients across all stages achieved a C-index of 0.739, outperforming a Cox regression model (0.716) and traditional TNM staging (0.706) for lung cancer-specific survival prediction. A separate DL neural network for surgically resected stage III NSCLC patients achieved a C-index of 0.843, substantially better than random forest (0.678) and Cox models (0.678).

Comprehensive AI models integrating multiple data types achieve the highest accuracy. An XGBoost model trained at Kyushu University combined 17 clinicopathological factors with 52 perioperative blood test values across stage IA to IIIA NSCLC patients. This model achieved AUC values of 0.890 for disease-free survival, 0.926 for overall survival, and 0.960 for cancer-specific survival - substantially outperforming pathological staging alone. The careful curation of multifaceted perioperative data was identified as the key driver of this high performance.

Genomic mutation models reach near-pathologist accuracy. A deep convolutional neural network developed by Coudray et al. trained on 1,634 patients from TCGA extracted mutation-related morphological features from hematoxylin and eosin histology images to predict the presence of specific genetic mutations (such as EGFR and KRAS) with AUC values up to 0.97 - substantially higher than previous models (0.75 to 0.83) and comparable to expert pathologist performance. This approach could enable mutation prediction directly from standard pathology slides without separate molecular testing.

TL;DR: AI approaches to NSCLC prediction range from machine learning on clinical data to deep learning on CT images and histology slides, with comprehensive multi-data models achieving AUC values above 0.90 - far surpassing traditional staging systems.
Pages 8-9
AI Models for Imaging and Treatment Response

CT imaging is a rich source of prognostic AI features. Radiomics models extract hundreds to thousands of quantitative features from CT scans capturing tumor texture, shape, and heterogeneity invisible to the human eye. Huang et al.'s radiomics signature for early-stage NSCLC achieved a C-index of 0.72 for disease-free survival - significantly better than conventional clinical factors (0.691). Zhang et al.'s Intelligent Prognosis Evaluation System analyzed CT images from 6,371 patients and achieved a C-index of 0.817, compared to 0.733 for Cox regression on clinical data alone.

3D convolutional neural networks process CT volumes effectively. Hosny et al.'s 3D CNN model trained on 1,194 NSCLC patients across multiple treatment cohorts achieved AUC values of 0.70 to 0.71 for predicting two-year overall survival in radiotherapy and surgery patients respectively, compared to just 0.55 to 0.58 for random forest models based on clinical information. The model identified large, uninterrupted areas of relatively higher CT density as key imaging predictors.

AI can predict immunotherapy response from pathology slides. The Deep-IO deep learning model was trained to predict response to immune checkpoint inhibitors from histopathology images, achieving an AUROC of 0.75 in internal testing and 0.66 in external validation. Importantly, combining Deep-IO with PD-L1 expression testing improved discriminative accuracy (AUROC 0.70) beyond either test alone, suggesting that AI pathology analysis complements - rather than replaces - standard biomarker testing for immunotherapy selection.

Integrating metabolomics with CT imaging pushes prediction further. A prospective study combining tissue metabolomics with CT imaging achieved a C-index of 0.74 (rising to 0.78 when combined with clinical data) - substantially outperforming radiomics alone (0.62) or CT alone (0.57). Key metabolites including sedoheptulose-7-phosphate and uridine showed high correlations with outcome, suggesting that metabolic profiling of tumor tissue represents a promising additional layer for NSCLC prediction models.

TL;DR: AI-driven imaging analysis using CT radiomics and 3D deep learning consistently outperforms traditional clinical staging for NSCLC prognosis, with multi-modal models combining imaging, pathology, and metabolomics achieving the highest predictive accuracy.
Pages 2, 5, 7
Key Prognostic Variables Across Clinical Models

Performance status is the most consistent clinical predictor. Across early-stage and metastatic NSCLC models, ECOG performance status - a standardized measure of how well a patient can carry out ordinary daily activities - emerged as a strong and consistent independent predictor of survival. Patients with ECOG PS 2 (limited self-care capacity) faced roughly double the mortality risk compared to those with PS 0 in immunotherapy settings.

Inflammatory markers reflect the immune landscape of disease. The neutrophil-to-lymphocyte ratio, lactate dehydrogenase (LDH), and the systemic immune-inflammatory index appear repeatedly across independent models as prognostic markers. These blood-based inflammatory indices capture the balance between pro-tumor and anti-tumor immune activity, and their elevation predicts worse outcomes whether patients receive chemotherapy, radiotherapy, or immunotherapy.

Tumor burden and metastatic site distribution matter significantly. In the metastatic setting, models consistently identified the number of metastatic sites as a strong predictor of survival and treatment response. Having three or more sites of metastasis was a particularly unfavorable prognostic marker in chemoimmunotherapy studies, while squamous histology conferred additional risk beyond staging in some analyses.

Molecular alterations guide targeted therapy selection and prognosis. EGFR mutation status and ALK gene rearrangements are the primary molecular predictors in NSCLC, directing the use of tyrosine kinase inhibitors. AI models trained on EGFR genotype combined with CT imaging (such as the Intelligent Prognosis Evaluation System) achieved superior survival prediction compared to clinical staging alone, suggesting that molecular and imaging data are complementary rather than redundant.

TL;DR: Across both conventional and AI-based NSCLC models, performance status, inflammatory blood markers, tumor burden, and molecular alterations like EGFR mutations emerge as consistently powerful prognostic variables that should be incorporated into future integrated models.
Pages 8-9
Barriers to Real-World Clinical Implementation

External validation is the most critical unmet need. The vast majority of prediction models reviewed were developed on retrospective, single-institution datasets and validated only in the same population or a closely related one. Without testing in geographically, ethnically, and clinically diverse external cohorts, it is impossible to know whether a model's performance will hold in the broader population of patients an oncologist actually treats.

Interpretability limits clinical trust in AI models. Deep learning models in particular achieve high predictive accuracy but function as 'black boxes' that cannot explain which features drove a specific prediction for a specific patient. Clinicians require not just accurate predictions but understandable ones - knowing why a model predicts high risk enables clinicians to confirm that the reasoning aligns with clinical experience. Models like XGBoost offer partial feature importance explanations, but fully transparent models with clinician-interpretable outputs remain underdeveloped.

Prospective trials are largely absent. Nearly all models in this review were developed retrospectively, using historical patient data to train and test algorithms. Prospective studies - in which the model is applied to new patients as they are treated, and outcomes are tracked forward in time - provide the most reliable evidence for clinical utility. Only one model in the AI table was developed using prospective data, highlighting a critical evidence gap.

Regulatory and guideline approval has not been granted to any model reviewed. None of the clinical or AI-based models surveyed had received formal guideline approval or regulatory endorsement as of the time of this review. All require further randomized controlled trial validation before they can be incorporated into standard clinical care pathways - a process that takes years and requires substantial institutional commitment.

TL;DR: Despite impressive performance metrics, no NSCLC prediction model reviewed has achieved guideline approval, held back by a lack of external validation, poor model interpretability, the near-complete absence of prospective trial data, and no regulatory endorsement.
Pages 7-9
Toward Personalized NSCLC Treatment Through Prediction

Prediction models could identify patients who need more aggressive treatment. A validated model for early-stage NSCLC that could identify aggressive stage I tumors likely to recur would enable oncologists to offer adjuvant chemotherapy to those patients - a therapy currently not offered routinely at that stage. Conversely, identifying low-risk patients who are unlikely to recur would spare them from toxic and expensive treatments with limited benefit.

Immunotherapy selection is a pressing clinical challenge that models could address. Not all NSCLC patients benefit from immune checkpoint inhibitors, and current biomarkers like PD-L1 expression are imperfect predictors of response. AI models that integrate imaging, genomic, and blood biomarker data could more precisely identify which patients will achieve durable benefit from immunotherapy - and which would be better served by alternative approaches - reducing both unnecessary toxicity and cost.

The convergence of large-scale databases and AI creates new opportunities. The growing availability of large pathology databases, multimodal tumor registries, and publicly accessible genomic datasets (like TCGA and GEO) provides the raw material for more robust model development and external validation. As institutions share data and adopt standardized collection protocols, the evidence base for NSCLC prediction models should improve substantially.

The ultimate goal is individualized treatment optimization. Rather than applying population-level guidelines uniformly, the vision articulated in this review is a future where each patient receives treatment recommendations tailored to their unique molecular profile, clinical characteristics, and imaging features. Prediction models - particularly AI-enhanced ones integrating all available data types - represent the most promising pathway toward that goal for patients with NSCLC.

TL;DR: Clinical and AI-based NSCLC prediction models have the potential to transform treatment selection by identifying patients who need more aggressive therapy, optimizing immunotherapy decisions, and ultimately enabling truly individualized treatment planning.
Citation: Open Access, 2025. Available at: PMC12529306.