The radiomics promise and its unfulfilled potential Radiomics - the extraction of hundreds of quantitative features from routine CT scans - has the potential to serve as a non-invasive biomarker for tumor biology, treatment response, and patient outcomes in lung cancer. Despite thousands of published studies, no radiomic biomarker has been widely accepted into clinical practice.
Why radiomic research often fails The lack of clinical translation is primarily due to poor research methodology: small cohorts, retrospective designs, inconsistent image acquisition protocols, non-standardized feature extraction, lack of external validation, and limited biological interpretation. Most studies do not follow the full radiomic workflow from study design through to critical appraisal.
Purpose of this guide Researchers from the University of Manchester and The Christie NHS Foundation Trust wrote this practical guide specifically for early career lung cancer researchers. It covers the complete radiomic workflow for CT-based lung cancer imaging - from study design, image acquisition, and segmentation through feature extraction, model building, validation, and critical appraisal.
Unique lung cancer challenges addressed Lung cancer CT imaging has specific challenges including breathing-induced tumor motion, atelectasis and emphysema obscuring tumor boundaries, varying reconstruction protocols, and differences between intravenous contrast and non-contrast scans. The guide addresses each of these challenges with practical solutions.
Traditional handcrafted features The conventional radiomics approach extracts hundreds to thousands of pre-defined features from a segmented region of interest. These fall into three categories: morphological/shape features (volume, surface area, sphericity), first-order/intensity features (mean, standard deviation, skewness of pixel intensities), and second-order/textural features (spatial patterns of intensity variation such as entropy, homogeneity, and correlation).
Deep learning featureless approaches An alternative approach uses convolutional neural networks to automatically learn features relevant to a specific clinical outcome from images, without manual feature engineering. These 'discovered' features may capture complex spatial patterns invisible to handcrafted descriptors, but require larger datasets and are less interpretable.
Advantages of radiomics over biopsies Radiomics is non-invasive, can analyze the entire tumor volume rather than a small biopsy sample, can capture intra-tumor heterogeneity, and can be performed on routinely acquired clinical scans without additional patient burden. It can also be repeated at multiple timepoints to track tumor evolution.
The reproducibility problem Despite these advantages, radiomic features are highly sensitive to differences in CT scanner type, acquisition parameters, reconstruction algorithms, and segmentation methods. A feature that appears predictive in one scanner or institution may not replicate at another, which is the fundamental challenge preventing clinical adoption.
Mandatory pre-study engagement Before conducting radiomic research, the guide recommends engaging with leading experts, early career researchers in similar fields, attending conferences, involving patient advocates (public patient involvement), and consulting statisticians for power calculations. Skipping these steps leads to underpowered studies with irrelevant questions.
Cohort definition pitfalls The most common design failures include heterogeneous cohorts (mixing different cancer stages, histologies, or treatment types), small sample sizes, and purely retrospective single-institution datasets. The guide recommends pre-specifying inclusion/exclusion criteria that reflect the real-world target population, and documenting any imaging exclusions.
Outcome variable selection The choice of clinical outcome must be clinically meaningful (e.g., overall survival, treatment response, toxicity risk) rather than chosen for statistical convenience. Confounding features - known clinical predictors of the outcome - must be collected and included in analyses. This ensures the radiomic biomarker adds value beyond existing clinical information.
Ethics, registration, and open science All radiomic studies should be registered on a clinical trials database before data collection begins. Regulatory approvals, data anonymization plans, and explicit data sharing commitments should be documented upfront. The RQS (Radiomic Quality Score) framework - a 16-point evaluation tool - should guide study design and serve as a self-assessment checklist.
Image acquisition variability Radiomic features are sensitive to CT acquisition parameters including slice thickness, pixel size, peak tube voltage, tube current, and use of intravenous contrast. Differences in scanner manufacturer and reconstruction algorithm (filtered back projection vs. iterative reconstruction) introduce systematic feature variations. Test-retest experiments - scanning the same patient twice - quantify this variability and help identify stable features.
Lung-specific motion challenges Free-breathing CT acquisitions used in routine lung cancer imaging capture tumor position at a random point in the breathing cycle, introducing inter-scan variability in feature values. 4D CT acquisitions capture multiple breathing phases but require specialized analysis. The guide recommends documenting the imaging phase and considering motion-robust features or image registration.
Segmentation as a source of error The region of interest (ROI) segmentation - manually contouring the tumor - is a major source of inter-observer variability in radiomics. Manual segmentation variability can cause the same tumor to generate different feature values depending on who drew the contour. Automated segmentation reduces this variability but introduces its own errors. The guide recommends reporting inter-observer agreement metrics.
IBSI standardization The Image Biomarker Standardisation Initiative (IBSI) provides internationally agreed nomenclature and definitions for radiomic features, benchmarking extraction algorithms across institutions. Researchers should use IBSI-compliant software and report compliance to ensure features have the same definition and interpretation across studies.
The overfitting problem Radiomics studies often extract hundreds of features from small patient cohorts. With more features than patients, any statistical model will overfit - appearing to predict well in training data but failing on new patients. Dimensionality reduction (e.g., removing redundant or unstable features), penalized regression (LASSO), and strict train/test splits are essential safeguards.
Feature selection requirements Only features demonstrated to be stable (reproducible across scanners and segmenters) should enter predictive models. Non-stable features add noise rather than signal. The guide recommends a staged selection: first assess stability, then assess predictive relevance only within the stable feature subset.
Validation hierarchy Internal validation (cross-validation, bootstrapping) is the minimum required, but is insufficient alone. External validation on a geographically and temporally distinct cohort acquired on different scanners is necessary to establish genuine generalizability. The validation cohort must be independent and pre-specified, not selected after initial model failure.
Clinical utility testing Statistical prediction alone is insufficient for clinical adoption. The radiomic model must be shown to improve upon existing clinical predictors (stage, histology, performance status) and provide decision-relevant information. Decision curve analysis tests whether acting on model predictions actually improves patient outcomes compared to treating all or no patients.
Current state of the field Despite thousands of radiomic publications, most have low quality scores on the RQS framework, lack external validation, and were never designed for clinical translation. The exponential increase in publications has not been matched by clinical progress. This credibility gap risks undermining the entire field.
Key recurring failures Table 2 in the guide summarizes the most common methodological failures: heterogeneous cohorts, small sample sizes, different scanners across institutions, ignoring respiratory motion, extracting large numbers of unstable features, no external validation, and limited biological interpretation linking imaging features to biology.
The delta radiomics opportunity Analyzing changes in radiomic features between scans at different timepoints (delta radiomics) can capture tumor response to treatment that static single-scan analysis misses. This approach is particularly relevant for monitoring early treatment response in lung cancer and may offer additional predictive value over baseline scans alone.
Future directions The field needs prospective multicenter studies with standardized imaging protocols and pre-specified endpoints. Integration of radiomics with clinical, pathological, and molecular (genomic) data is needed to understand biological mechanisms and compare radiomic value to existing predictors. Federated learning approaches that train models across institutions without sharing patient data may overcome the data sharing barriers that limit large multicenter validation.