Predictive Value of Machine Learning for Radiation Pneumonitis and Checkpoint Inhibitor Pneumonitis in Lung Cancer Patients: A Systematic Review and Meta-Analysis

Sci Rep 2025 AI 5 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
The Growing Problem of Treatment-Related Pneumonitis in Lung Cancer

Two Types of Treatment-Related Lung Toxicity Lung cancer patients face pneumonitis risk from two increasingly common treatment modalities. Radiation pneumonitis (RP) occurs in approximately 15-40% of patients receiving thoracic radiotherapy, caused by radiation-induced inflammatory injury to lung parenchyma. Checkpoint inhibitor pneumonitis (CIP) affects 3-10% of patients receiving immune checkpoint inhibitors, caused by T-cell-mediated autoimmune lung injury. Both can be life-threatening and force treatment interruption.

Why Prediction Matters If high-risk patients could be identified before treatment, oncologists could modify radiation dose constraints, use alternative immunotherapy agents, increase monitoring frequency, or initiate prophylactic interventions. Currently, neither RP nor CIP can be reliably predicted using standard clinical factors, creating a clear unmet need that machine learning models could address.

The Machine Learning Landscape Over the past decade, numerous studies have developed ML models for RP and CIP prediction using diverse feature sets - clinical variables, dosimetric parameters from radiation plans, radiomic imaging features, and combinations thereof. However, the field lacks synthesis: no comprehensive review had quantified overall predictive performance across all ML approaches.

Study Scope This systematic review and meta-analysis identified 56 eligible studies comprising 12,803 patients. Studies were analyzed separately for RP and CIP prediction, with subgroup analyses by feature type (clinical only, dosimetric only, radiomic only, or multimodal combinations). Quality was assessed using the PROBAST framework and radiomic quality was scored using the Radiomic Quality Score (RQS).

TL;DR: This systematic review synthesized 56 studies and 12,803 patients to quantify overall machine learning performance for predicting radiation pneumonitis and checkpoint inhibitor pneumonitis in lung cancer, finding that multimodal models outperform single-modality approaches.
Pages 2-3
PROBAST Assessment and Radiomic Quality Scoring

PROBAST Framework for Bias Assessment The Prediction model Risk Of Bias Assessment Tool (PROBAST) was used to assess methodological quality across four domains: participant selection, predictors, outcome, and statistical analysis. Each study was rated as low, high, or unclear risk of bias. This systematic quality evaluation allows readers to understand which studies' results can be most trusted.

Radiomic Quality Score (RQS) For studies using radiomic features, the Radiomic Quality Score - a 36-point checklist covering data acquisition, segmentation, feature extraction, model development, and reporting - was applied. Higher RQS indicates better methodological rigor. A median RQS of 12/36 across the included radiomic studies reveals that the field has significant methodological gaps to address.

Meta-Analysis Methodology Concordance statistics (c-index) were pooled using random-effects meta-analysis to account for heterogeneity between studies. Separate pooled estimates were calculated for RP and CIP prediction, and for each feature combination type. Heterogeneity was quantified using the I-squared statistic, and funnel plot asymmetry was examined for publication bias.

Feature Type Classification Studies were classified by the primary feature types used: clinical variables (patient characteristics, performance status, lung function), dosimetric parameters (mean lung dose, V20, V5), radiomic features extracted from CT images, and dosiomics features (radiomic-like analysis applied to dose distribution maps). Combinations of these feature types were also analyzed separately.

TL;DR: PROBAST assessed methodological quality and a median Radiomic Quality Score of 12/36 revealed significant gaps in radiomic study rigor, while meta-analysis pooled concordance statistics across feature type subgroups.
Pages 3-5
Multimodal Models Best Predict Radiation Pneumonitis

Pooled Performance by Feature Type For RP prediction, clinical-only models achieved pooled c-index around 0.70. Dosimetric-only models performed similarly (c-index ~0.72). Radiomic-only models showed higher performance (c-index ~0.78). The strongest performance came from combined radiomics-plus-dosiomics models, which achieved pooled c-index of 0.90 - a substantial and clinically meaningful improvement over single-modality approaches.

Dosiomics as an Emerging Feature Type Dosiomics refers to applying radiomic-style texture analysis to radiation dose distribution maps rather than anatomical images. Unlike mean lung dose or V20 which summarize dose in simple scalar metrics, dosiomics captures spatial dose heterogeneity patterns. The high performance of radiomics-dosiomics combinations suggests that spatial information from both anatomy and dose distribution provides complementary predictive signal.

Volume Effect Matters Models incorporating dosimetric variables consistently outperformed clinical-only models, confirming that the volume of lung receiving high radiation dose (captured by metrics like V20 and mean lung dose) is a stronger RP predictor than clinical patient characteristics alone. Combining anatomical radiomic heterogeneity with dosimetric spatial patterns achieves the best predictions.

High Heterogeneity Between Studies I-squared statistics were high for all subgroup analyses, indicating substantial variability between studies beyond random chance. This likely reflects differences in patient populations, radiation techniques, RP grading systems used, and ML algorithm choice. High heterogeneity means pooled estimates should be interpreted cautiously.

TL;DR: Combined radiomics-dosiomics models achieved the highest pooled c-index of 0.90 for radiation pneumonitis prediction, substantially outperforming clinical-only (0.70), dosimetric-only (0.72), or radiomic-only (0.78) approaches.
Pages 5-6
Radiomic-Clinical Combinations Lead for CIP Prediction

Smaller CIP Evidence Base Fewer studies were available for CIP prediction compared to RP (the field is newer, given that checkpoint inhibitors were only approved in the mid-2010s). Despite this, a consistent pattern emerged: models combining radiomic imaging features with clinical variables achieved the highest performance, with pooled c-index of 0.86.

The Radiomic Signal for CIP Baseline CT radiomic features of normal-appearing lung parenchyma predicted subsequent CIP development. Features capturing ground-glass texture and density heterogeneity in pre-treatment scans were particularly informative. This suggests that subclinical lung vulnerability - potentially reflecting underlying interstitial changes - is detectable on CT before immunotherapy begins.

Clinical Predictors of CIP Among clinical variables, prior thoracic radiation history, pre-existing interstitial lung disease, and baseline performance status emerged as consistent CIP predictors across multiple studies. The combination of these clinical risk factors with radiomic lung texture features produced the strongest models.

Immunotherapy Specifics Some studies distinguished between anti-PD-1, anti-PD-L1, and combination checkpoint inhibitor regimens in their CIP analyses. Combination checkpoint inhibitor therapy (anti-PD-1 plus anti-CTLA-4) showed higher CIP rates, and models built specifically for combination regimens generally outperformed general models, suggesting that CIP prediction may need to be agent-specific.

TL;DR: Radiomic-clinical combined models achieved pooled c-index of 0.86 for CIP prediction, with baseline CT lung texture features capturing subclinical lung vulnerability that predicts immunotherapy-related pneumonitis risk.
Pages 6-7
Raising the Quality Bar in Pneumonitis Prediction Research

Low Radiomic Quality Scores The median RQS of 12/36 across radiomic studies indicates widespread methodological deficiencies. Common problems included lack of test-retest reproducibility analysis, absence of external validation, failure to report intraclass correlation coefficients for inter-reader segmentation reliability, and inadequate description of CT acquisition parameters. Future studies must raise their methodological standards.

Prospective Study Design Needed Nearly all included studies were retrospective. Prospective studies with pre-specified endpoints, standardized CT acquisition protocols, and centralized radiomic processing would substantially improve the reliability of model estimates and reduce risk of overfitting.

Standardization of RP and CIP Grading Different studies used different grading systems for RP (CTCAE, RTOG) and different definitions of clinically significant versus subclinical pneumonitis. Standardized, prospectively defined endpoints are essential for comparing models across studies and combining data in meta-analyses.

Clinical Implementation Gap High model performance in retrospective studies does not guarantee clinical utility. Prospective clinical decision-support trials, where the model's output is provided to treating clinicians and its impact on treatment decisions and patient outcomes is measured, are needed before these models can be recommended for routine clinical use.

TL;DR: The field must address low radiomic quality scores, move toward prospective study designs, standardize pneumonitis grading definitions, and ultimately conduct clinical utility trials before ML pneumonitis models can be deployed in routine cancer care.
Citation: Open Access, 2025. Available at: PMC12214896.