Targeted Therapy Requires Molecular Profiling. Non-small cell lung cancer (NSCLC) accounts for 80-85% of all lung cancers. Molecular subtyping has become essential for genotype-driven targeted therapy, which is now the standard of care for many patients with advanced NSCLC. Three key driver mutations -- EGFR, ALK, and KRAS -- determine eligibility for specific targeted agents that can dramatically improve outcomes.
Limitations of Current Testing Methods. Conventional molecular genotyping requires invasive tissue biopsies and genetic sequencing, facing challenges including high costs, sampling bias, inadequate samples (too few cells or hemorrhagic samples), lengthy turnaround times, procedural risks, and medical complications. Liquid biopsy from blood offers an alternative but remains limited in accessibility and cost for many patients.
Radiomics as a Non-Invasive Alternative. Radiomics uses machine learning to extract and analyze quantitative features from medical images -- typically CT scans -- to model tumor characteristics without tissue sampling. The combination of radiomics and AI has shown promise as a non-invasive tool for predicting oncogene mutation status in NSCLC, using imaging data already acquired as part of standard clinical care.
Scope of This Analysis. This PRISMA-compliant systematic review and meta-analysis evaluated the effectiveness of radiomics alone or combined with clinical data for predicting EGFR, ALK, and KRAS mutation status in NSCLC. It assessed AI model performance, whether clinical data improves radiomics-based prediction, and the clinical implications of current evidence.
Search Strategy and Inclusion. A systematic search through Medline, Cochrane Library, and EMBASE identified 890 articles, which after de-duplication and screening yielded 124 studies included in the qualitative analysis and 51 eligible for meta-analysis. Studies covered AI-based models predicting EGFR, ALK, and KRAS mutation status in NSCLC patients using CT-derived radiomics.
Types of AI Methods Used. Of the 124 included studies, 90 exclusively used machine learning algorithms, most commonly logistic regression (56 studies), support vector machine (49 studies), and random forest (39 studies). Only 10 studies exclusively used deep learning algorithms, and 10 used classical statistical models. The most common validation approach was training-validation split rather than external validation, which was performed in only 21 studies.
Imaging and Segmentation Approaches. CT was the most frequently used imaging modality (86 studies), followed by PET/CT (30 studies). Manual tumor segmentation was used in 73 studies, while automatic or semi-automatic methods were used in 42. Most studies used non-contrast-enhanced CT images. The lack of standardization across these methodological choices is a major source of variability in reported results.
Patient Population. The 124 studies encompassed 45,747 NSCLC patients with a median age of approximately 62 years. Most patient datasets originated from Chinese cohorts (87 of 124 studies), followed by the United States (9 studies). The majority of studies focused on adenocarcinoma histology and included treatment-naive patients imaged before any therapeutic intervention, though TNM stage was often not reported or varied widely across studies.
EGFR Mutation Prediction. The meta-analysis of 37 studies (72 different models) predicting EGFR mutation status from CT-derived radiomics demonstrated an overall AUC of 0.769. The pooled sensitivity was 0.754 (95% CI: 0.727-0.780), meaning the AI correctly identified approximately 75% of EGFR-mutated tumors. The false positive rate was 0.344 (95% CI: 0.308-0.381), corresponding to a pooled specificity of 0.656.
EGFR with Clinical Data Added. When clinical variables were combined with radiomics features, a separate meta-analysis showed improved performance for EGFR mutation prediction, with sensitivity rising to 0.806 (95% CI: 0.777-0.833) and the false positive rate decreasing to 0.315 (95% CI: 0.270-0.364). The AUC increased from 0.769 (radiomics only) to 0.821 (combined model) -- a modest but potentially meaningful 5% improvement.
ALK Translocation Prediction. The meta-analysis for ALK translocation included three studies (four models) and showed an overall AUC of 0.831 with a pooled sensitivity of 0.754 (95% CI: 0.638-0.841). The false positive rate was 0.225 (95% CI: 0.163-0.302), giving a specificity of 0.775. A positive ALK result was 11 times more likely to be detected than missed (diagnostic odds ratio 11.10).
Substantial Study Heterogeneity. Individual study sensitivities for EGFR ranged from 0.196 to 0.982, and false positive rates ranged from 0.006 to 0.761 -- reflecting enormous variability across studies. This heterogeneity is a critical finding: it means that results from any individual study may not generalize to other clinical settings, and pooled estimates must be interpreted cautiously.
KRAS Mutation Prediction. Six studies (eight models) evaluated radiomics-based prediction of KRAS mutations. The overall AUC was 0.735, but the pooled sensitivity was only 0.475 (95% CI: 0.153-0.820) -- far lower than for EGFR or ALK. Individual study sensitivities ranged from 0.015 to 0.875, indicating extreme variability. A positive KRAS prediction was only about 4 times more likely to be correct than incorrect (diagnostic odds ratio 4.28).
Why KRAS Performance Is Inconsistent. The poor aggregate KRAS sensitivity was heavily influenced by two models from a single study that recorded sensitivity near zero. This single study dramatically skewed the pooled estimate. The extreme heterogeneity in KRAS prediction results means current evidence is exploratory rather than conclusive, and standardization of modeling strategies is urgently needed before any clinical application.
Meta-Regression Findings. A meta-regression analyzing factors that might explain variability in EGFR prediction performance found no statistically significant influence from age, use of contrast enhancement, segmentation type (manual vs. semi-automatic), model type, or AI methodology (machine learning vs. deep learning). This null result likely reflects the limited number of studies and substantial missing data rather than a genuine absence of influence.
Publication Bias Assessment. Deeks' funnel plot asymmetry tests found no significant publication bias for EGFR (radiomics-only), ALK, or KRAS predictions. A statistically significant result emerged only for combined EGFR models, but sensitivity analysis showed this was driven by a single influential study. These findings support the integrity of the included literature, though the evidence base remains limited in depth for ALK and KRAS.
A Complement, Not a Replacement. Radiomics-based AI models are not intended to replace gold standard molecular testing such as PCR or next-generation sequencing. Rather, they could serve as non-invasive pre-screening tools to optimize pre-test probability, helping identify patients most likely to carry targetable mutations and prioritizing them for laboratory-based confirmation. This approach could save time, reduce costs, and preserve valuable invasively-obtained tissue samples.
The False Positive Challenge. Like low-dose CT screening for lung cancer, AI-based mutation prediction tools carry the risk of false positives, which could lead to unnecessary confirmatory tests or misinterpretations. Strategies to mitigate this include multimodal integration of clinical and radiological data to enhance specificity, stricter decision thresholds in screening contexts, and hierarchical workflows combining complementary predictive tools.
False Negatives Are Equally Important. While false positives create unnecessary downstream workup, false negatives may cause missed opportunities for targeted therapy in patients with actionable mutations. For mutation-screening tools, sensitivity is the priority metric: a missed mutation in a patient who would benefit from a targeted agent is a serious clinical failure that directly impacts survival.
The Added Value of Clinical Variables. Although adding clinical data to radiomics models showed only a modest 5% AUC improvement for EGFR prediction, this gain may have meaningful implications in large-scale screening contexts where even small improvements in discriminative ability translate to substantial differences in patient outcomes and healthcare resource allocation. Clinical variables may especially help reduce false positives and enhance specificity in triage workflows.
Segmentation Variability. Interobserver variability in manual tumor segmentation significantly affects the ability of radiomics to predict oncogene mutations. Automatic or semi-automatic segmentation approaches may be preferable for reducing this variability. The lack of segmentation standardization across studies is one reason model performance varies so dramatically even within the same mutation type.
Feature Harmonization Across Centers. Scanner and protocol differences between institutions can compromise model generalizability. One study applying ComBat harmonization to radiomics features from multicenter CT datasets reported a 10-15% improvement in predictive performance for both EGFR and KRAS mutation prediction. This underscores that multicenter models require explicit harmonization techniques to be clinically viable.
Machine Learning vs. Deep Learning. No statistically significant differences in EGFR prediction performance were found between traditional machine learning and deep learning approaches. However, deep learning models -- particularly end-to-end convolutional neural networks -- offer a practical advantage: they require only a lesion bounding box rather than full segmentation, reducing the variability introduced by the segmentation step while potentially capturing complex spatial patterns not accessible to hand-crafted radiomics features.
CT Protocol Standardization. CT slice thickness and experimental settings variability also affect radiomics-based model predictiveness. Studies have shown that these technical parameters influence which features can be reliably extracted and compared across datasets. Future studies should use standardized, IBSI (Image Biomarker Standardisation Initiative)-compliant feature extraction to enable meaningful cross-study comparison.
Critical Evidence Gaps for ALK and KRAS. Only three studies were available for ALK meta-analysis and six for KRAS, making it impossible to draw firm conclusions for these mutations. Future studies specifically targeting ALK and KRAS prediction with adequate sample sizes and multicenter designs are urgently needed to determine whether radiomics tools can reliably identify these mutations non-invasively.
Sample Size and External Validation. Over half of the 124 included studies had sample sizes below 200 patients, and only 21 studies used external cohorts for validation. For AI-based models in medical imaging, both large sample sizes and independent external validation are essential to demonstrate that performance generalizes beyond the development dataset. These are the most critical gaps in the current literature.
Prioritize Advanced-Stage Patients. Future studies should focus on stage III-IV NSCLC patients -- particularly stage IV -- where molecular testing is clinically mandatory per guidelines. Most existing studies did not report TNM stage or included patients across all stages, diluting clinical relevance. Models developed and validated in the exact population where mutation testing is most needed would have the greatest translational value.
Conclusion. Radiomics-based AI models demonstrate meaningful predictive ability for EGFR and ALK mutation status in NSCLC, with EGFR models showing the strongest evidence base. KRAS prediction remains inconsistent and requires further standardization. As non-invasive triage tools to optimize molecular testing workflows, these approaches show real promise -- but robust prospective validation and methodological standardization are prerequisites for clinical translation.