Lung cancer is the leading cause of cancer death worldwide, accounting for roughly one in five cancer-related deaths. The central problem is timing: most cases are not caught until they are at an advanced stage, when curative options are limited. Low-dose CT (LDCT) scanning is the gold standard for early screening, but a single LDCT study involves hundreds of image slices, and large screening programs generate thousands of scans per week.
This volume creates practical problems: reader fatigue among radiologists, inter-observer variability (two different radiologists interpreting the same scan differently), and time constraints that can cause subtle findings to be missed. AI and machine learning offer the promise of consistent, automated first-pass analysis that can prioritize which scans need urgent attention and flag suspicious findings that might otherwise be overlooked.
This narrative review searched PubMed, Scopus, IEEE Xplore, and Google Scholar for studies published between 2018 and 2025 using AI or machine learning for lung cancer imaging, detection, and prognosis. After removing duplicates and applying inclusion criteria, 89 studies were included: 33 review articles and 58 original research papers. Studies were categorized into three themes: detection and screening, risk prediction and prognosis, and staging and diagnosis.
Study quality was assessed using established frameworks including PROBAST and QUADAS-2. The review specifically prioritized studies that reported external validation, multi-center testing, or clinically meaningful performance metrics. Commonly used benchmark datasets included LIDC-IDRI, LUNA16, and NLST - publicly available lung CT scan databases that allow comparison between different AI models.
Pulmonary nodules - small rounded opacities in the lung - are the earliest imaging sign of lung cancer. Several deep learning models have achieved impressive performance in detecting these nodules on LDCT. NoduleX, a hybrid model combining CNN features with radiomics, achieved an AUC of 0.99 for malignancy discrimination on the LIDC-IDRI dataset. ResNet-50-based models trained on tens of thousands of Chinese screening CT scans achieved AUC of 0.90 on independent data. More recent attention-based 3D models achieve sensitivity above 93% with as low as one false positive per scan.
Reducing false positives - cases where the AI incorrectly flags normal structures as suspicious - is critical for clinical adoption. Excessive false alarms erode radiologist confidence and lead to unnecessary follow-up procedures. Modern multi-stage detection systems use an initial high-sensitivity detector followed by a specialized false-positive reduction module. These second-stage modules are trained to recognize common benign mimics like vascular branching points, airway walls, and pleural thickenings that can resemble nodules to earlier-generation models.
Beyond detection, AI is being used to predict how patients will do over time - overall survival, disease recurrence, and treatment response. Radiomics pipelines extract hundreds of quantitative features from CT or PET/CT images: measurements of nodule shape, texture, density, and internal heterogeneity that are invisible to the naked eye. When combined with clinical variables like age, smoking history, and tumor stage, radiomics models consistently outperform staging-only predictions.
Deep survival models - neural networks adapted for time-to-event data - and their classical counterparts like penalized Cox regression and random survival forests are both used. A key finding across multiple studies is that simple multimodal fusion (combining imaging features with a small set of clinical variables like age, stage, and histology) often delivers most of the performance benefit without requiring complex multi-modal architectures. Longitudinal approaches that track changes between serial CT scans further improve prediction by capturing tumor growth dynamics.
Accurate TNM staging - determining tumor size and whether cancer has spread to lymph nodes or distant organs - is essential for treatment planning but is difficult even for experienced radiologists, particularly for lymph node assessment. AI algorithms applied to CT and PET-CT imaging can improve non-small-cell lung cancer (NSCLC) staging accuracy, lymph node evaluation, and malignancy subtype classification, frequently matching or exceeding radiologist performance.
Deep learning models for segmentation - delineating exactly where a tumor ends and normal tissue begins - are also improving. U-Net architectures and their 3D variants achieve Dice scores above 0.80 for nodules larger than 10 mm, providing precise volumetric measurements needed for treatment monitoring and surgical planning. Attention-based models are particularly effective for small or irregular nodules where boundary detection is most challenging.
Despite strong performance on benchmark datasets, most AI lung cancer tools have not been adopted in routine clinical care. The core problem is generalizability: a model trained on one institution's scans - with its specific CT scanner model, reconstruction settings, patient demographics, and radiologist annotation conventions - often performs worse when applied to another institution's data. This phenomenon, called domain shift, means that high AUC values reported in studies frequently overstate real-world performance.
External validation using truly independent datasets from different hospitals and countries remains rare. Most studies use internal splits of a single dataset, which does not adequately test generalizability. Additionally, many AI models function as 'black boxes' - they produce a prediction without explaining which features drove it. Clinicians are rightly cautious about relying on recommendations they cannot understand or verify. Explainable AI techniques that highlight which image regions or features influenced a prediction are an important emerging area that addresses this trust gap.
The review concludes that AI and machine learning have genuine transformative potential for lung cancer care - in detection, staging, and prognosis - but must clear important hurdles before delivering that promise in routine clinical practice. Priority areas include: standardizing imaging protocols so that data from different hospitals is more comparable, creating large multi-institutional datasets that include diverse racial and geographic populations, and conducting prospective clinical trials where AI outputs actually influence clinical decisions and patient outcomes are tracked.
Regulatory approval, ethical AI implementation, and clinician training are also identified as necessary components of successful translation. AI tools that triage worklists, assist with difficult nodule characterization, and provide calibrated risk estimates have the clearest near-term pathway to clinical adoption when paired with appropriate human oversight. The review frames AI not as a replacement for expert clinical judgment but as a powerful tool to standardize, scale, and enhance it.