Ensemble Machine Learning Classifiers Combining CT Radiomics and Clinical-Radiological Features for Preoperative Prediction of Pathological Invasiveness in Lung Adenocarcinoma Presenting as Part-Solid Nodules: A Multicenter Retrospective Study

Technol Cancer Res Treat 2025 AI 6 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Predicting Invasiveness of Part-Solid Lung Adenocarcinoma with Ensemble ML

Clinical Problem: Part-solid nodules in lung adenocarcinoma span a spectrum from adenocarcinoma in situ (AIS) and minimally invasive adenocarcinoma (MIA) to fully invasive adenocarcinoma (IAC), each with different surgical requirements and prognosis.

Study Innovation: This multicenter retrospective study is the first to systematically compare ensemble machine learning classifiers combining CT radiomics and clinical-radiological features for preoperative invasiveness prediction in part-solid nodule adenocarcinoma.

Study Scale: 344 patients from 3 centers were included, with 1,239 radiomic features extracted and LASSO regression used for feature selection, followed by training of multiple ensemble ML architectures.

Best Performance: The stacking ensemble classifier achieved AUC 0.84, accuracy 0.817, and F1 score 0.869, outperforming individual classifiers and providing clinically actionable invasiveness predictions.

TL;DR: A multicenter study of 344 patients found that ensemble ML classifiers combining CT radiomics with clinical features predict adenocarcinoma invasiveness with AUC 0.84 and F1 0.869 in part-solid nodules.
Pages 2-3
Part-Solid Nodules and Surgical Planning in Lung Adenocarcinoma

Part-Solid Nodule Spectrum: Part-solid nodules (PSNs) with a ground-glass opacity component represent a histological spectrum from pre-invasive to invasive lung adenocarcinoma. The solid component size correlates with invasiveness.

Treatment Implications: AIS and MIA can typically be managed with sublobar resection (wedge or segmentectomy), preserving more lung function, while IAC generally requires lobectomy to achieve adequate oncological margins.

Preoperative Challenge: Current LDCT-based risk stratification tools inadequately distinguish AIS/MIA from IAC, resulting in either over-treatment (lobectomy for potentially non-invasive disease) or under-treatment (sublobar resection for IAC).

Radiomics Opportunity: CT radiomics can extract quantitative features from the solid component, ground-glass component, and their interface that reflect the underlying histological invasiveness grade non-invasively.

TL;DR: Distinguishing AIS/MIA from IAC in part-solid nodules preoperatively determines the optimal surgical extent; CT radiomics can non-invasively capture imaging correlates of histological invasiveness.
Pages 3-4
Radiomics Pipeline and Ensemble ML Architecture

Feature Extraction: 1,239 radiomic features were extracted using PyRadiomics from manually segmented nodule ROIs on thin-slice CT, including first-order, shape, GLCM, GLRLM, and GLSZM texture features.

LASSO Feature Selection: LASSO regression with 10-fold cross-validation reduced 1,239 features to a compact predictive subset, controlling for multicollinearity and overfitting in the limited-sample setting.

Ensemble Architectures Tested: Multiple ensemble strategies were compared: bagging, boosting (gradient boosting, XGBoost), random forest, and stacking - which combines predictions from multiple base classifiers using a meta-learner.

Clinical-Radiological Integration: Clinical features (age, sex, smoking) and radiological features (nodule size, consolidation tumor ratio, CT morphological signs) were combined with the radiomics score in the final models.

TL;DR: 1,239 radiomic features were extracted, LASSO-selected, and used to train four ensemble ML architectures combined with clinical-radiological features for invasiveness prediction.
Pages 5-6
Stacking Classifier Outperforms Individual Models

Stacking Best Performance: The stacking classifier - which uses a meta-learner to optimally combine predictions from multiple base classifiers - achieved AUC 0.84, accuracy 0.817, and F1 0.869, outperforming all individual algorithms.

F1 Score Emphasis: The high F1 score (0.869) indicates that the model maintains both high precision (few false IAC predictions) and high recall (few missed IAC cases), a balanced performance relevant to the clinical binary decision.

Comparison to Single Classifiers: Standalone classifiers (logistic regression, single decision tree, SVM) showed lower and more variable performance, confirming the advantage of ensemble combination in this heterogeneous multicenter dataset.

Multicenter Validation: Performance was validated across all three participating centers, with consistent results indicating that the model generalizes beyond single-institution data - critical for intended clinical deployment.

TL;DR: Stacking ensemble achieved the best AUC (0.84), accuracy (0.817), and F1 (0.869), consistently outperforming individual classifiers across all three centers.
Pages 7-8
Surgical Planning and Treatment Personalization

Surgical Extent Guidance: A preoperative invasiveness prediction tool could guide surgeons toward sublobar resection for predicted AIS/MIA and recommend lobectomy for predicted IAC, personalizing the surgical approach based on AI-assisted risk stratification.

Shared Decision-Making: The nomogram output could be presented to patients during preoperative consultation to facilitate informed discussion about the potential extent of resection and the rationale for each surgical option.

Intraoperative Relevance: In centers where intraoperative frozen section analysis is limited or unavailable, preoperative AI-based invasiveness prediction provides an additional layer of decision support for the operating surgeon.

Quality of Life Impact: Appropriately expanding sublobar resection for true AIS/MIA patients preserves lung function and improves long-term quality of life compared to lobectomy, making accurate preoperative classification directly patient-beneficial.

TL;DR: The stacking classifier could guide surgical extent decisions - directing sublobar resection for predicted AIS/MIA - improving quality of life without compromising oncological outcomes.
Pages 9-10
Study Limitations and Paths Forward

Retrospective Design: As a retrospective study, the dataset reflects past clinical practices including CT acquisition protocols that may differ from current standards, potentially affecting radiomics feature reproducibility.

Manual Segmentation: Manual ROI delineation of part-solid nodules - particularly the boundary between solid and ground-glass components - introduces variability. Automated deep learning-based segmentation would improve reproducibility.

Generalizability Questions: Despite multicenter validation across 3 sites, external validation in geographically diverse cohorts with different scanner brands and reconstruction algorithms is needed to confirm broad applicability.

Prospective Integration: A prospective study comparing AI-guided versus standard surgical planning for PSN patients - measuring surgical extent decisions, pathological concordance, and patient outcomes - would definitively establish clinical utility.

TL;DR: Future work requires automated PSN segmentation, expanded multicenter external validation, and prospective surgical outcome studies to establish clinical utility of the stacking ensemble model.
Citation: Open Access, 2025. Available at: PMC12174711.