Global burden of lung cancer Lung cancer is one of the deadliest cancers worldwide, responsible for approximately 18% of all cancer-related deaths. Cigarette smoking remains the primary risk factor, though the disease affects a wide range of patients.
Diagnostic complexity Accurate classification of lung cancer requires microscopic evaluation of tissue specimens combined with immunohistochemical staining. For the roughly 70% of patients with advanced, unresectable disease, only small biopsies or cytology specimens are available, making precise interpretation both critical and challenging.
Role of AI in pathology Deep learning, particularly convolutional neural networks (CNNs), is increasingly applied to digital pathology. These methods can assist pathologists by automating subtyping, reducing inter-observer variability, and preserving limited tissue samples for molecular testing.
Scope of this review This systematic review covered 96 studies identified from a PubMed search through March 2023, evaluating deep learning applications for lung cancer diagnosis, classification, prognosis, mutational status characterization, and PD-L1 expression estimation from histological and cytological images.
Search strategy The authors searched PubMed using terms combining deep learning, CNNs, and lung cancer with histological or cytological image analysis. The search covered all publications from inception through March 2023.
Inclusion and exclusion criteria Studies were included if they developed at least one deep learning model for histopathological or cytopathological assessment of malignant lung lesions. Reviews, editorials, non-English articles, and non-human studies were excluded.
Study selection process Four authors independently screened citations using the Rayyan platform. Full texts of potentially eligible articles were reviewed, with disagreements resolved by consensus. The final dataset included 96 studies.
Data extraction From each paper, researchers extracted the first author, publication year, medical aim, technical method, classification problem, dataset details, and performance metrics. This structured approach enabled direct comparisons across studies.
Binary tumor detection Several studies trained CNNs to distinguish cancerous from normal lung tissue. Models such as EfficientNet-B3 trained with weakly supervised learning achieved AUC values of 0.97-0.98 across multiple independent validation datasets.
Main histological subtypes The most common classification task was distinguishing adenocarcinoma (ADC), squamous cell carcinoma (SCC), and small cell lung carcinoma (SCLC). Multi-class models using architectures such as ResNet-152 and VGG-19 achieved macro-average AUCs above 0.90.
Rare neuroendocrine tumors A subset of studies tackled rarer subtypes including large cell neuroendocrine carcinoma (LCNEC) and atypical carcinoid. One model using HALO-AI achieved 98% accuracy distinguishing SCLC, LCNEC, and atypical carcinoid from 150 whole slide images.
Immunohistochemical feature prediction An innovative system (WIFPS) predicted IHC biomarker expression directly from H and E-stained slides, achieving Cohen's kappa values of 0.79-0.89 in validation cohorts for three-class ADC/SCC/SCLC classification.
Survival prediction from whole slide images Multiple studies developed deep learning models to predict overall survival and recurrence risk directly from H and E-stained slides, without requiring additional molecular testing.
Adenocarcinoma architectural patterns A key prognostic task involved classifying the dominant growth pattern of lung adenocarcinomas (lepidic, acinar, papillary, micropapillary, solid), which correlates strongly with patient outcomes. DL models showed high concordance with expert pathologists.
Multi-modal integration Some studies combined image-based features with clinical or genomic data to improve survival prediction. Models integrating pathology images with clinical variables consistently outperformed single-modality approaches.
Practical implications Automated prognostic stratification from routine pathology slides could allow risk-adapted treatment planning without additional costly tests, potentially benefiting patients where molecular profiling is not available.
EGFR and other driver mutations Studies evaluated DL models for predicting EGFR, KRAS, ALK, and RET mutation status from H and E slides. Performance was variable - EGFR prediction showed moderate AUC values, while ALK status prediction achieved AUC above 0.90 in some cohorts.
PD-L1 expression estimation PD-L1 tumor proportion score (TPS) determines eligibility for first-line immunotherapy. Deep learning models attempted to estimate PD-L1 positivity directly from H and E slides, though results were more modest compared to diagnosis tasks.
Cytological specimens Some studies extended DL methods to cytology samples such as bronchial washings and fine needle aspirates. These models achieved high accuracy despite the more variable image quality compared to histological sections.
Clinical value If validated prospectively, image-based molecular prediction could prioritize which samples require expensive molecular testing, preserving limited biopsy material and accelerating treatment decisions for patients with advanced disease.
Addressing workforce shortages Pathologist shortages worldwide create delays in diagnosis. AI-assisted analysis could reduce the time from specimen receipt to report, particularly for straightforward classification tasks.
Reducing subjectivity Deep learning models provide reproducible, quantitative assessments that eliminate inter-observer variability in subjective tasks such as growth pattern scoring and PD-L1 estimation.
Preserving diagnostic material For patients with limited biopsy samples, AI-guided prioritization of staining and testing could reduce unnecessary tissue consumption while ensuring comprehensive diagnosis.
Regulatory and workflow considerations Deployment of DL models in clinical settings requires regulatory approval, integration with laboratory information systems, and workflow redesign. Most reviewed models were developed on retrospective datasets and have not undergone prospective clinical validation.
Dataset limitations Most studies used relatively small, single-institution datasets. External validation on diverse patient populations and different staining protocols is essential before clinical implementation.
Explainability gap Many high-performing CNN models function as black boxes, making it difficult to understand the morphological features driving predictions. Explainability tools and pathologist-AI collaboration frameworks are needed.
Standardization challenges Variability in tissue preparation, staining, and digital scanning creates domain shift problems. Models trained at one institution may not generalize to another without additional adaptation.
Future research priorities Prospective multicenter trials, development of federated learning approaches to protect patient data, and integration of pathology AI with genomics and radiology represent the key directions for advancing this field toward real-world clinical impact.