EGFR mutations are the most therapeutically actionable genetic alterations in lung adenocarcinoma, the most prevalent subtype of non-small cell lung cancer, which accounts for 80 to 85 percent of all lung cancers. Patients with EGFR-mutant tumors respond dramatically to EGFR tyrosine kinase inhibitors (TKIs), making accurate EGFR status determination a prerequisite for treatment selection.
The current gold standard for EGFR testing is molecular analysis of tumor tissue obtained through needle biopsy or surgical resection. However, biopsy is invasive, procedurally risky in patients with centrally located or poorly accessible tumors, and may yield false-negative results due to intratumoral heterogeneity where sampled tissue may not represent the full mutational landscape of the tumor.
18F-FDG PET/CT is already acquired routinely in lung cancer staging. EGFR-mutant cells have altered glucose metabolism regulated by EGFR signaling, creating an imaging phenotype potentially distinguishable from wild-type tumors. Extracting this information quantitatively through radiomics and deep learning would enable noninvasive EGFR status prediction from an already-acquired scan without additional procedures.
Prior studies have predicted EGFR status using clinical variables alone, or combined clinical and radiomic features, achieving AUCs of approximately 0.80. This study develops and validates a three-component Clinical-Radiomics-Deep Learning (CRD) model integrating all available information from PET/CT, aiming to improve performance beyond what any single-modality or two-component approach achieves.
218 patients with histologically confirmed lung adenocarcinoma who underwent 18F-FDG PET/CT before any treatment were retrospectively analyzed, with 122 classified as EGFR-mutant and 96 as EGFR wild-type. The dataset was randomly split 7:3 into training (n=153) and test (n=65) cohorts. All patients received EGFR mutation testing through real-time fluorescence PCR on histological samples, targeting mutations in exons 18 through 21.
PET/CT acquisition followed Image Biomarker Standardization Initiative guidelines using a standardized protocol on the same scanner system across all patients. Patients fasted 6 to 8 hours before imaging, and 18F-FDG was administered at 3.7 MBq/kg with image acquisition beginning 60 minutes postinjection. CT parameters included 120 kV, 100 mA, and 3 mm slice thickness.
Quantitative metabolic parameters were computed from PET images including SUVmax, SUVmean, metabolic tumor volume (MTV) using the adaptive threshold method, and total lesion glycolysis (TLG) calculated as SUVmean multiplied by MTV. Volumes of interest were manually adjusted by experienced radiologists in three planes to ensure full tumor coverage.
3D tumor segmentation was performed by two nuclear medicine physicians with over 10 years of experience using 3D Slicer, with a semiautomatic 40% SUVmax threshold applied and inter-observer agreement required to exceed 0.75 overlap coefficient. Discordant cases were adjudicated by a senior third radiologist.
Three independent feature sets were extracted: clinical variables, radiomic features from PET/CT images, and deep learning features from five neural network architectures. Clinical features included age, sex, smoking history, TNM stage, SUVmax, MTV, and TLG. Univariate and multivariate logistic regression identified SUVmax as the single stable clinical predictor retained for model construction.
Radiomic feature extraction used PyRadiomics after 3D wavelet transformation to generate eight sub-band images. From the original and transformed images, 851 features were derived spanning shape, first-order statistics, GLCM texture, GLSZM, GLRLM, GLDM, and NGTDM feature families. Z-score normalization followed by t-test filtering and LASSO regression with 10-fold cross-validation reduced this to a stable radiomic signature.
Deep learning features were extracted from 2D ROI images of the largest PET/CT slice for each lesion, sized to 224x224 pixels, from five independently trained architectures: DenseNet169 (1,664 features), ResNet50 (2,048 features), ConvNext (1,024 features), Swin Transformer (1,024 features), and EfficientNet (1,280 features). Features were taken from the penultimate layer of each trained network.
Feature selection used 5-fold cross-validation within the training set, retaining features selected in at least three of five folds. Stable LASSO coefficients averaged across folds produced a radiomic score and a deep learning score per patient, which together with the clinical feature formed the inputs for the three comparative models. ConvNext features combined with logistic regression achieved the highest validation AUC of 0.890 and were selected as the deep learning component.
The CRD model achieved an AUC of 0.821 on the independent test set, significantly outperforming both the clinical-only model (C model, AUC 0.599) and the combined clinical-radiomics model (CR model, AUC 0.739) by DeLong test. The CRD model's advantage over the C model was highly significant (Z = -3.522, p less than 0.001, corrected p = 0.001), and its advantage over the CR model was also statistically significant (Z = -2.197, corrected p = 0.028).
On the training cohort, the performance hierarchy was similarly consistent: CRD model AUC 0.919, CR model AUC 0.851, and C model AUC 0.660, indicating that addition of first radiomics and then deep learning features progressively improved predictive power without overfitting to training data.
An ablation study confirmed the contribution of each feature type. The radiomics-deep learning model (RD, without clinical data) achieved an AUC of 0.815. The clinical-deep learning model (CD, without radiomics) achieved an AUC of 0.821, matching the full CRD model on AUC but showing inferior PPV and NPV. The full CRD model achieved the highest PPV (0.800) and NPV (0.700) among all model combinations, providing the most reliable predictions in both mutation-positive and mutation-negative calls.
SHAP analysis quantified the contribution of each component to CRD predictions. Deep learning features showed the most consistent directional impact, with higher feature values concentrated in positive SHAP regions indicating strong promotion of EGFR-mutant prediction. Radiomic features showed more balanced contributions, while clinical features (SUVmax) showed dispersed SHAP values reflecting their variable impact across different patient profiles.
A nomogram was constructed from the CRD model to enable individualized EGFR mutation probability estimation at the point of care. The nomogram integrates three inputs: the clinical score (SUVmax), the radiomic score, and the deep learning score from ConvNext features, assigning weighted points to each that sum to a total score convertible to predicted EGFR mutation probability.
Calibration curve analysis showed good agreement between predicted probabilities from the CRD nomogram and observed EGFR mutation rates in both training and test cohorts, indicating that the model's probability estimates are accurate and not systematically over- or under-predicting risk. Calibration was assessed using 1,000 bootstrap resamples.
Decision curve analysis (DCA) demonstrated that the CRD model provided greater net clinical benefit across the full range of threshold probabilities compared to both the C and CR models, as well as compared to the treat-all and treat-none reference strategies. This confirms that acting on CRD model predictions would improve clinical decision-making across a wide range of acceptable risk thresholds.
The balanced sensitivity (0.757) and specificity (0.750) of the CRD model on the test set reflect its clinical practicality: it avoids the common trade-off where high sensitivity for mutation detection comes at the cost of high false-positive rates that would lead to unnecessary TKI treatment in wild-type patients.
The progression from AUC 0.599 (clinical alone) to 0.739 (clinical plus radiomics) to 0.821 (clinical plus radiomics plus deep learning) demonstrates that each feature type captures distinct and complementary information about EGFR genotype. Clinical variables reflect population-level epidemiologic associations such as female sex and nonsmoking status that are well-established EGFR mutation correlates. Radiomics captures macro-level tumor heterogeneity through quantitative texture metrics.
Deep learning adds a third dimension: automatically learned hierarchical representations that capture both the tumor and its microenvironment, including pleural traction features associated with EGFR-mutant adenocarcinomas, without requiring labor-intensive manual feature specification. The model also benefits from capturing information from the surrounding tissue rather than only within the segmented tumor volume.
ConvNext emerged as the superior deep learning architecture, outperforming DenseNet169, ResNet50, EfficientNet, and Swin Transformer. ConvNext is a modern convolutional architecture that matches transformer performance while maintaining the computational efficiency of CNNs, and its superior performance on this task suggests that local convolutional feature patterns are more relevant to EGFR genotype detection than global attention-based representations in this imaging context.
Limitations include exclusive use of Asian patient cohorts from a single institution, restricting generalizability to other ethnic groups where EGFR mutation prevalence and tumor biology differ. The 2D single-slice deep learning approach, while computationally efficient, does not capture full 3D volumetric context. Rare EGFR mutations outside exons 18 to 21 were excluded to maintain group homogeneity for the primary classification task.
The CRD model's most immediate clinical application is as a noninvasive first-pass screening tool to identify likely EGFR-mutant patients from their pretreatment PET/CT, which is already routinely acquired for staging. For patients where biopsy yields wild-type EGFR results, the model could serve as a complementary check, flagging cases where imaging phenotype strongly suggests mutation that may have been missed due to sampling bias from intratumoral heterogeneity.
For patients where biopsy is contraindicated due to tumor location, comorbidities, or patient refusal, the CRD model could provide clinically actionable genotypic information that might otherwise be unavailable, potentially allowing initiation of TKI therapy based on imaging evidence in carefully considered cases.
The nomogram format translates the model's prediction into an individualized probability score that clinicians can discuss with patients, supporting shared decision-making about whether to proceed with repeat biopsy, initiate empirical TKI therapy, or pursue alternative diagnostic workups. Decision curve analysis confirms net clinical benefit across a wide probability threshold range, meaning the model adds value even when different clinicians set different risk thresholds for action.
Future work planned by the authors includes prospective multicenter validation with diverse ethnic cohorts, extension to predict ALK, ROS-1, and other actionable mutations, development of 3D volumetric deep learning models as larger datasets become available, and integration of explainable AI methods to improve transparency and regulatory compliance for clinical deployment.
This study demonstrates that integrating clinical variables, quantitative radiomic features, and automatically learned deep learning representations from 18F-FDG PET/CT produces a significantly more accurate EGFR mutation prediction model than any single data source or two-source combination. The AUC improvement from 0.599 to 0.821 reflects the additive value of each informational layer.
The ConvNext-based deep learning component was the single most powerful predictor, confirming that modern convolutional architectures applied to PET/CT imaging slices can learn EGFR-relevant imaging phenotypes that neither clinical data nor handcrafted radiomic features fully capture.
The nomogram and decision curve framework translate the model into a clinically usable tool that is compatible with standard oncology workflows, requiring only PET/CT image input and basic clinical variables that are routinely documented for every lung cancer patient at staging.
Validated prospectively in larger, multicenter, and multi-ethnic cohorts, the CRD framework could establish noninvasive imaging-based EGFR genotyping as a standard component of lung adenocarcinoma workup, complementing tissue-based testing and expanding precision oncology access to patients for whom biopsy is not feasible.