Prediction of oncogene mutation status in non-small cell lung cancer: a systematic review and meta-analysis with a special focus on artificial intelligence-based methods

Eur Radiol 2026 AI 7 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Why Non-Invasive Mutation Testing Matters

Targeted Therapy Requires Molecular Profiling. Non-small cell lung cancer (NSCLC) accounts for 80-85% of all lung cancers. Molecular subtyping has become essential for genotype-driven targeted therapy, which is now the standard of care for many patients with advanced NSCLC. Three key driver mutations -- EGFR, ALK, and KRAS -- determine eligibility for specific targeted agents that can dramatically improve outcomes.

Limitations of Current Testing Methods. Conventional molecular genotyping requires invasive tissue biopsies and genetic sequencing, facing challenges including high costs, sampling bias, inadequate samples (too few cells or hemorrhagic samples), lengthy turnaround times, procedural risks, and medical complications. Liquid biopsy from blood offers an alternative but remains limited in accessibility and cost for many patients.

Radiomics as a Non-Invasive Alternative. Radiomics uses machine learning to extract and analyze quantitative features from medical images -- typically CT scans -- to model tumor characteristics without tissue sampling. The combination of radiomics and AI has shown promise as a non-invasive tool for predicting oncogene mutation status in NSCLC, using imaging data already acquired as part of standard clinical care.

Scope of This Analysis. This PRISMA-compliant systematic review and meta-analysis evaluated the effectiveness of radiomics alone or combined with clinical data for predicting EGFR, ALK, and KRAS mutation status in NSCLC. It assessed AI model performance, whether clinical data improves radiomics-based prediction, and the clinical implications of current evidence.

TL;DR: Non-invasive prediction of NSCLC driver mutations using CT-based radiomics and AI could supplement or reduce the need for invasive biopsies, critical for enabling timely targeted therapy decisions.
Pages 3, 4, 11, 12
Systematic Review Design and Study Landscape

Search Strategy and Inclusion. A systematic search through Medline, Cochrane Library, and EMBASE identified 890 articles, which after de-duplication and screening yielded 124 studies included in the qualitative analysis and 51 eligible for meta-analysis. Studies covered AI-based models predicting EGFR, ALK, and KRAS mutation status in NSCLC patients using CT-derived radiomics.

Types of AI Methods Used. Of the 124 included studies, 90 exclusively used machine learning algorithms, most commonly logistic regression (56 studies), support vector machine (49 studies), and random forest (39 studies). Only 10 studies exclusively used deep learning algorithms, and 10 used classical statistical models. The most common validation approach was training-validation split rather than external validation, which was performed in only 21 studies.

Imaging and Segmentation Approaches. CT was the most frequently used imaging modality (86 studies), followed by PET/CT (30 studies). Manual tumor segmentation was used in 73 studies, while automatic or semi-automatic methods were used in 42. Most studies used non-contrast-enhanced CT images. The lack of standardization across these methodological choices is a major source of variability in reported results.

Patient Population. The 124 studies encompassed 45,747 NSCLC patients with a median age of approximately 62 years. Most patient datasets originated from Chinese cohorts (87 of 124 studies), followed by the United States (9 studies). The majority of studies focused on adenocarcinoma histology and included treatment-naive patients imaged before any therapeutic intervention, though TNM stage was often not reported or varied widely across studies.

TL;DR: The systematic review included 124 studies covering 45,747 NSCLC patients, revealing widespread reliance on single-center machine learning approaches with limited external validation, predominantly from Chinese cohorts studying adenocarcinoma.
Pages 12, 18
AI Performance for EGFR and ALK Prediction

EGFR Mutation Prediction. The meta-analysis of 37 studies (72 different models) predicting EGFR mutation status from CT-derived radiomics demonstrated an overall AUC of 0.769. The pooled sensitivity was 0.754 (95% CI: 0.727-0.780), meaning the AI correctly identified approximately 75% of EGFR-mutated tumors. The false positive rate was 0.344 (95% CI: 0.308-0.381), corresponding to a pooled specificity of 0.656.

EGFR with Clinical Data Added. When clinical variables were combined with radiomics features, a separate meta-analysis showed improved performance for EGFR mutation prediction, with sensitivity rising to 0.806 (95% CI: 0.777-0.833) and the false positive rate decreasing to 0.315 (95% CI: 0.270-0.364). The AUC increased from 0.769 (radiomics only) to 0.821 (combined model) -- a modest but potentially meaningful 5% improvement.

ALK Translocation Prediction. The meta-analysis for ALK translocation included three studies (four models) and showed an overall AUC of 0.831 with a pooled sensitivity of 0.754 (95% CI: 0.638-0.841). The false positive rate was 0.225 (95% CI: 0.163-0.302), giving a specificity of 0.775. A positive ALK result was 11 times more likely to be detected than missed (diagnostic odds ratio 11.10).

Substantial Study Heterogeneity. Individual study sensitivities for EGFR ranged from 0.196 to 0.982, and false positive rates ranged from 0.006 to 0.761 -- reflecting enormous variability across studies. This heterogeneity is a critical finding: it means that results from any individual study may not generalize to other clinical settings, and pooled estimates must be interpreted cautiously.

TL;DR: CT-based radiomics AI achieved AUC values of 0.769 for EGFR and 0.831 for ALK mutation prediction, with combined radiomics plus clinical data models modestly improving EGFR performance to AUC 0.821.
Pages 18-19
KRAS Prediction and Meta-Regression Findings

KRAS Mutation Prediction. Six studies (eight models) evaluated radiomics-based prediction of KRAS mutations. The overall AUC was 0.735, but the pooled sensitivity was only 0.475 (95% CI: 0.153-0.820) -- far lower than for EGFR or ALK. Individual study sensitivities ranged from 0.015 to 0.875, indicating extreme variability. A positive KRAS prediction was only about 4 times more likely to be correct than incorrect (diagnostic odds ratio 4.28).

Why KRAS Performance Is Inconsistent. The poor aggregate KRAS sensitivity was heavily influenced by two models from a single study that recorded sensitivity near zero. This single study dramatically skewed the pooled estimate. The extreme heterogeneity in KRAS prediction results means current evidence is exploratory rather than conclusive, and standardization of modeling strategies is urgently needed before any clinical application.

Meta-Regression Findings. A meta-regression analyzing factors that might explain variability in EGFR prediction performance found no statistically significant influence from age, use of contrast enhancement, segmentation type (manual vs. semi-automatic), model type, or AI methodology (machine learning vs. deep learning). This null result likely reflects the limited number of studies and substantial missing data rather than a genuine absence of influence.

Publication Bias Assessment. Deeks' funnel plot asymmetry tests found no significant publication bias for EGFR (radiomics-only), ALK, or KRAS predictions. A statistically significant result emerged only for combined EGFR models, but sensitivity analysis showed this was driven by a single influential study. These findings support the integrity of the included literature, though the evidence base remains limited in depth for ALK and KRAS.

TL;DR: KRAS mutation prediction remains unreliable with current radiomics approaches, showing extreme inter-study variability and low pooled sensitivity, while meta-regression identified no significant predictors of model performance.
Pages 20-23
Clinical Implications of Radiomics-Based Mutation Testing

A Complement, Not a Replacement. Radiomics-based AI models are not intended to replace gold standard molecular testing such as PCR or next-generation sequencing. Rather, they could serve as non-invasive pre-screening tools to optimize pre-test probability, helping identify patients most likely to carry targetable mutations and prioritizing them for laboratory-based confirmation. This approach could save time, reduce costs, and preserve valuable invasively-obtained tissue samples.

The False Positive Challenge. Like low-dose CT screening for lung cancer, AI-based mutation prediction tools carry the risk of false positives, which could lead to unnecessary confirmatory tests or misinterpretations. Strategies to mitigate this include multimodal integration of clinical and radiological data to enhance specificity, stricter decision thresholds in screening contexts, and hierarchical workflows combining complementary predictive tools.

False Negatives Are Equally Important. While false positives create unnecessary downstream workup, false negatives may cause missed opportunities for targeted therapy in patients with actionable mutations. For mutation-screening tools, sensitivity is the priority metric: a missed mutation in a patient who would benefit from a targeted agent is a serious clinical failure that directly impacts survival.

The Added Value of Clinical Variables. Although adding clinical data to radiomics models showed only a modest 5% AUC improvement for EGFR prediction, this gain may have meaningful implications in large-scale screening contexts where even small improvements in discriminative ability translate to substantial differences in patient outcomes and healthcare resource allocation. Clinical variables may especially help reduce false positives and enhance specificity in triage workflows.

TL;DR: Radiomics AI is best positioned as a pre-screening triage tool to identify mutation-likely candidates for molecular testing, not as a replacement for biopsy-based molecular profiling in clinical decision-making.
Pages 23-24
Methodological Factors Affecting Model Quality

Segmentation Variability. Interobserver variability in manual tumor segmentation significantly affects the ability of radiomics to predict oncogene mutations. Automatic or semi-automatic segmentation approaches may be preferable for reducing this variability. The lack of segmentation standardization across studies is one reason model performance varies so dramatically even within the same mutation type.

Feature Harmonization Across Centers. Scanner and protocol differences between institutions can compromise model generalizability. One study applying ComBat harmonization to radiomics features from multicenter CT datasets reported a 10-15% improvement in predictive performance for both EGFR and KRAS mutation prediction. This underscores that multicenter models require explicit harmonization techniques to be clinically viable.

Machine Learning vs. Deep Learning. No statistically significant differences in EGFR prediction performance were found between traditional machine learning and deep learning approaches. However, deep learning models -- particularly end-to-end convolutional neural networks -- offer a practical advantage: they require only a lesion bounding box rather than full segmentation, reducing the variability introduced by the segmentation step while potentially capturing complex spatial patterns not accessible to hand-crafted radiomics features.

CT Protocol Standardization. CT slice thickness and experimental settings variability also affect radiomics-based model predictiveness. Studies have shown that these technical parameters influence which features can be reliably extracted and compared across datasets. Future studies should use standardized, IBSI (Image Biomarker Standardisation Initiative)-compliant feature extraction to enable meaningful cross-study comparison.

TL;DR: Segmentation variability, scanner heterogeneity, and CT protocol differences are major sources of performance variation in radiomics models; harmonization techniques and standardized protocols are critical for multicenter deployment.
Pages 23-24
Future Research Priorities and Evidence Gaps

Critical Evidence Gaps for ALK and KRAS. Only three studies were available for ALK meta-analysis and six for KRAS, making it impossible to draw firm conclusions for these mutations. Future studies specifically targeting ALK and KRAS prediction with adequate sample sizes and multicenter designs are urgently needed to determine whether radiomics tools can reliably identify these mutations non-invasively.

Sample Size and External Validation. Over half of the 124 included studies had sample sizes below 200 patients, and only 21 studies used external cohorts for validation. For AI-based models in medical imaging, both large sample sizes and independent external validation are essential to demonstrate that performance generalizes beyond the development dataset. These are the most critical gaps in the current literature.

Prioritize Advanced-Stage Patients. Future studies should focus on stage III-IV NSCLC patients -- particularly stage IV -- where molecular testing is clinically mandatory per guidelines. Most existing studies did not report TNM stage or included patients across all stages, diluting clinical relevance. Models developed and validated in the exact population where mutation testing is most needed would have the greatest translational value.

Conclusion. Radiomics-based AI models demonstrate meaningful predictive ability for EGFR and ALK mutation status in NSCLC, with EGFR models showing the strongest evidence base. KRAS prediction remains inconsistent and requires further standardization. As non-invasive triage tools to optimize molecular testing workflows, these approaches show real promise -- but robust prospective validation and methodological standardization are prerequisites for clinical translation.

TL;DR: Radiomics AI shows promising evidence for EGFR and ALK mutation prediction in NSCLC, but clinical translation requires multicenter prospective validation, larger sample sizes, and standardized imaging and modeling protocols.
Citation: Open Access, 2026. Available at: PMC12963223.