Digital Biomarkers for Precision Early Detection of Lung Cancer: Integrating AI-Driven Multi-Omics Into Clinical Pathways

Cancer Med 2026 AI 7 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Why Lung Cancer Demands Better Early Detection Tools

The mortality burden. Lung cancer is the leading cause of cancer-related death worldwide, responsible for more deaths than breast, colorectal, and prostate cancers combined. The fundamental reason for this mortality is that most patients are diagnosed at late stages, when curative options are limited. Early-stage detection, before cancer has spread, is the single most impactful opportunity to improve survival.

Limitations of low-dose CT screening. Low-dose computed tomography (LDCT) is the only screening method with proven mortality benefit in high-risk populations, but it faces persistent limitations. High false-positive rates lead to unnecessary follow-up procedures, patient anxiety, and healthcare costs. A substantial proportion of detected nodules are indeterminate, requiring invasive biopsy or repeated scans to resolve. These challenges restrict widespread clinical adoption of LDCT alone as a population screening tool.

The multi-omics opportunity. Modern molecular profiling technologies - genomics, epigenomics, transcriptomics, proteomics, metabolomics, and microbiomics - can detect molecular signatures of early lung cancer in blood, sputum, breath, and urine before CT scans reveal visible tumors. Each omics layer captures a different dimension of the biological changes occurring as normal cells transform into cancer, and integrating these layers provides a more complete and accurate picture than any single test.

AI as the integration engine. The data generated by multi-omics profiling is too large, complex, and high-dimensional for human analysis alone. Artificial intelligence - particularly machine learning and deep learning - provides the computational framework to extract diagnostic signals from these massive datasets, identify the combinations of markers most predictive of early cancer, and embed these insights into clinical decision support tools. This review systematically evaluates both the biomarkers and the AI methods driving this field.

TL;DR: LDCT alone cannot solve early lung cancer detection; multi-omics technologies offer molecular biomarkers detectable years before imaging findings, and AI provides the framework to integrate these data into clinical decision support.
Pages 2-4
Genomic and Epigenomic Biomarkers

Driver gene mutations as early signals. Lung cancer develops through sequential accumulation of genetic alterations in oncogenes and tumor suppressor genes. Driver mutations in EGFR, KRAS, BRAF, HER2, and ALK are detectable in circulating tumor DNA (ctDNA) from blood, sputum, and bronchoalveolar lavage fluid before tumors are clinically apparent. Multi-target panels combining TP53 sequencing, KRAS mutation analysis, p16 methylation, and microsatellite instability assessment detect tumor-associated alterations with differential sensitivity: TP53 mutations at 56%, KRAS at 27%, and p16 methylation at 38% in bronchoalveolar lavage samples.

DNA methylation as a sensitive early marker. Aberrant methylation of tumor suppressor gene promoters - particularly hypermethylation at CpG islands - occurs early in lung carcinogenesis, before invasive cancer forms. A multigene methylation panel targeting TAC1, HOXA17, and SOX17 in sputum achieved 98% sensitivity and 71% specificity across 150 NSCLC patients and 60 healthy controls. Critically, CDKN2A methylation is detectable in sputum samples collected up to three years before lung cancer diagnosis, providing a meaningful window for preventive intervention.

Multi-gene panels outperform single markers. A methylation panel targeting RASSF1A, SHOX2, and PTGER4 achieved AUC improvement from 0.69 to 0.74 when combined versus used individually. Another panel covering SOX17, HOXA9, AJAP1, PTGDR, UNCX, and MARCH11 reached 96.7% sensitivity and 60% specificity, while PCDHGB6, HOXA9, MGMT, and miR-126 together achieved 85.2% sensitivity and 81.5% specificity. These improvements confirm that combining markers from the same omics layer already substantially enhances detection accuracy.

Histology-specific epigenetic biomarkers. Meta-analysis of DNA methylation across histological subtypes of NSCLC reveals subtype-specific patterns. CDH13 and APC hypermethylation are more prevalent in adenocarcinoma versus squamous cell carcinoma (AUC 0.68 and 0.66 respectively), suggesting that methylation panels can not only detect lung cancer but also provide non-invasive guidance on tumor subtype - information that influences treatment selection.

TL;DR: Multi-gene DNA methylation panels in sputum and blood detect lung cancer up to 3 years before imaging, with sensitivity up to 98%, outperforming single-marker approaches and capturing histology-specific signals.
Pages 4-5
Transcriptomic and Non-Coding RNA Biomarkers

MicroRNA panels in blood. MicroRNAs are small non-coding RNAs that regulate gene expression, and their levels in blood reflect the activity of tumor cells. A five-miRNA panel (miR-20a, miR-223, miR-21, miR-221, miR-145) showed individual AUC values of 0.89 to 0.94 for NSCLC detection. Distinct miRNA profiles differ between adenocarcinoma and squamous cell carcinoma: miR-944 shows AUC 0.982 for squamous carcinoma while miR-3662 achieves AUC 0.926 for adenocarcinoma, enabling both detection and subtype classification from blood.

Multi-miRNA serum panels with exceptional performance. A four-miRNA serum panel developed and externally validated achieved an AUC of 0.993, among the highest reported for any blood-based lung cancer test. Exosomal miRNA panels - measuring small RNAs packaged in tumor-derived vesicles circulating in blood - achieved AUC values of 0.899 for NSCLC and 0.936 for adenocarcinoma specifically. Three miRNA classifiers from plasma samples discriminated SCLC from NSCLC with AUC 0.878 in training and 0.869 in validation, supporting histological subtyping from blood alone.

Long non-coding RNAs and circular RNAs. Long non-coding RNAs (lncRNAs) have tissue-specific expression patterns that improve diagnostic specificity. Plasma HOTAIR levels correlate with NSCLC progression and metastasis. The lncRNA GAS5 combined with standard protein markers CEA and CA199 achieved AUC 0.734 for early-stage NSCLC detection. Circular RNAs (circRNAs) are structurally resistant to degradation, making them unusually stable in blood and exosomes. Specific circRNAs including hsa_circ_0077837 and hsa_circ_0001821 distinguish NSCLC from normal tissue, and a meta-analysis of circRNAs in Chinese patients reported a pooled AUC of 0.78 for lung cancer diagnosis.

Airway epithelial gene expression classifiers. RNA from bronchial epithelial cells captured during bronchoscopy reflects field cancerization - the molecular changes spread throughout the airway in smokers with cancer. A 17-gene classifier built from 232 cancer-associated transcripts in bronchial cells from 299 smokers achieved AUC 0.78 for detecting lung cancer in patients with nondiagnostic bronchoscopy and AUC 0.81 in an independent cohort, with a negative predictive value of 94%. This approach is now clinically available as the Percepta Genomic Sequencing Classifier.

TL;DR: Blood-based miRNA panels achieve AUC up to 0.993 for NSCLC detection and can differentiate subtypes; airway epithelial gene expression classifiers provide negative predictive values near 94% to safely rule out cancer after nondiagnostic bronchoscopy.
Pages 5-6
Protein, Metabolite, Volatile, and Microbiome Biomarkers

Serum protein panels. Multiple protein markers show diagnostic utility for lung cancer, particularly when combined. CYFRA 21-1 is the strongest single predictor for NSCLC (AUC 0.78 to 0.847), while ProGRP is most sensitive for SCLC (AUC 0.875 to 0.86). Combining CYFRA 21-1, CEA, ProGRP, and NSE significantly improves SCLC diagnostic accuracy beyond any individual marker. Glycoprotein-enriched PON1 combined with AACT achieved AUC 0.940 with 94.4% sensitivity and 90.2% specificity for early NSCLC, approaching the performance needed for population screening.

Metabolic signatures in plasma and urine. Cancer cells exhibit fundamentally altered metabolism including enhanced glycolysis (the Warburg effect) and dysregulated lipid metabolism. A seven-phospholipid panel derived from plasma LC-MS analysis yielded AUC 0.88 for early-stage lung cancer detection. A non-targeted lipidomics study identified nine plasma phospholipids achieving 100% specificity in independent validation, with 90% or greater sensitivity in a cohort of 1,036 LDCT screening participants. Elevated urinary creatine and creatinine correlate with lung cancer risk across European and non-European populations and are detectable in serum and saliva as well.

Exhaled breath volatile organic compounds. Tumor cells release distinctive volatile organic compounds (VOCs) into exhaled breath that can be measured non-invasively. A six-VOC panel (including 2-butanone, 3-hydroxy-2-butanone, 4-hydroxyhexenal, acrolein, and malondialdehyde) achieved sensitivity of 96% or greater and specificity up to 100% in non-smokers. Detection of three or more elevated VOC markers achieves 0.95 specificity for distinguishing lung cancer from healthy controls, making breath testing potentially the least invasive detection modality.

Microbiome signatures. The lung, gut, and blood microbiome composition differs significantly between lung cancer patients and healthy individuals. An airway microbial signature combining Veillonella and Megasphaera achieved AUC 0.88 with 95% sensitivity and 75% specificity. A blood plasma microbiome model reached AUC 0.95 (sensitivity 0.81, specificity 0.90) and maintained robust performance (AUC 0.93 to 0.921) in two independent validation cohorts. While mechanism and causality remain under investigation, microbiome-based markers provide orthogonal diagnostic information not captured by molecular tumor profiling.

TL;DR: Protein panels, plasma phospholipid metabolites, exhaled breath VOCs, and microbiome signatures each provide complementary diagnostic windows into early lung cancer, with multiple approaches reaching clinically meaningful AUC values of 0.88 to 0.95.
Pages 8-9
AI and Machine Learning for Multi-Omics Integration

Supervised learning for classification. Machine learning models trained on labeled molecular data enable detection of lung cancer and histological subtyping as classification tasks. Convolutional neural networks applied to CT-derived radiomic features learn tumor shape, texture, and intensity patterns that map to molecular subtypes and clinical outcomes. An ensemble of five deep learning classifiers for lung adenocarcinoma achieved 99.2% prediction accuracy (AUC 0.988). A genomic deep learning model trained on whole-exome sequencing from 12 cancer types achieved AUC 0.94 for distinguishing cancer from normal tissue.

Unsupervised learning for subtype discovery. Variational autoencoders compress high-dimensional genomic or epigenomic profiles into compact latent representations that reveal molecularly distinct patient subgroups. A variational autoencoder trained on DNA methylation data to differentiate LUAD from LUSC achieved near-perfect discrimination (AUC approximately 1.0), demonstrating that epigenetic features alone can define the molecular identity of lung cancer subtypes. Unsupervised approaches are particularly valuable for discovering novel patient subgroups not defined by prior clinical categories.

Multi-modal fusion architectures. The most powerful AI frameworks combine multiple omics layers simultaneously using architectures designed for cross-modal data. Self-attention-based deep learning networks learn joint latent representations that capture inter-omics relationships, outperforming simple concatenation of single-omics features. Graph neural networks combine multi-omics data with graph autoencoder or graph attention mechanisms to model complex interactions among features across molecular layers. A deep neural network leveraging Kullback-Leibler divergence and focal loss for early cancer prediction achieved AUC 0.99, outperforming traditional methods.

AI-driven histopathology. Convolutional neural networks applied to whole-slide histopathology images from TCGA accurately distinguish LUAD from LUSC and link quantitative image features to transcriptomic subtypes (p less than 0.01). An AI algorithm classifying lung cancer slides achieves AUC 0.97 for subtype classification while simultaneously predicting driver gene mutations including STK11, EGFR, KRAS, and TP53 directly from tissue morphology. Transfer learning across over 17,355 slides from 28 tumor types demonstrates that computational histopathology features associate with genomic alterations, immune infiltration, and patient prognosis.

TL;DR: Deep learning models integrating multi-omics data achieve AUC values near 1.0 for cancer detection and subtype classification; fusion architectures that combine molecular and imaging data capture complementary signals unavailable to single-modality approaches.
Pages 9, 10, 12
Clinical Decision Support and Translation to Practice

AI as a complement to LDCT, not a replacement. Multi-omics AI tools are best positioned as decision support tools that work alongside LDCT rather than replacing it. The Percepta Genomic Sequencing Classifier uses airway gene expression to re-stratify cancer risk after nondiagnostic bronchoscopy, enabling both down-classification of low-risk nodules to avoid unnecessary procedures and escalation of high-risk cases for earlier biopsy. The Galleri multi-cancer early detection test integrates cell-free DNA methylation patterns and has demonstrated high specificity in large validation cohorts as an adjunct to LDCT in risk refinement.

Longitudinal monitoring and dynamic risk updating. A key advantage of liquid biopsy multi-omics approaches is their ability to track molecular changes over time. Longitudinal integration of ctDNA or methylation signals into AI-driven surveillance algorithms enables dynamic risk updating in high-risk individuals or patients with stable nodules, providing continuous monitoring that static imaging cannot offer. Prospective multicenter studies including CCGA and PATHFINDER are evaluating this real-world feasibility.

Explainable AI for clinical trust. State-of-the-art deep learning models are often criticized as opaque black boxes. Explainable AI (XAI) methods address this: Grad-CAM generates spatial heatmaps showing which regions of CT images drive model decisions, allowing clinicians to verify that the model focuses on medically relevant structures. SHAP values quantify which molecular features contribute most to classification, linking model predictions to biologically meaningful determinants. LIME builds locally interpretable approximations of complex model behavior for individual patient cases. These XAI approaches are essential for regulatory approval and clinical adoption.

Federated learning for multi-center collaboration. A critical barrier to AI model development is that patient data cannot be freely shared across institutions due to privacy regulations. Federated learning overcomes this by training AI models collaboratively across multiple institutions without exchanging raw patient data - each site trains on its own data and shares only model parameters. Applied to lung cancer cohorts, this approach accelerates biomarker validation at scale while maintaining compliance with patient privacy requirements.

TL;DR: Multi-omics AI tools function as embedded clinical decision support companions to LDCT, with explainable AI improving clinician trust, and federated learning enabling privacy-preserving multi-center model development.
Pages 10-14
Challenges, Emerging Technologies, and Future Directions

Data heterogeneity and integration barriers. Multi-omics data from different platforms - sequencing machines, mass spectrometers, microarrays - differ in formats, scales, and quality in ways that complicate integration. Inadequate batch effect correction can introduce technical artifacts that mislead AI models. Standardized preprocessing protocols, harmonized data formats, and robust batch correction methods are prerequisites for reliable multi-center model training and validation.

Annotation scarcity and demographic gaps. Current studies rely predominantly on single-institution cohorts with limited racial and demographic diversity. Public datasets frequently contain incomplete clinical annotations and are biased toward advanced-stage cancers, underrepresenting the early-stage cases most relevant for screening tool development. Creating large-scale, multiethnic, early-stage lung cancer repositories with comprehensive longitudinal clinical data is a critical infrastructure need for the field.

Single-cell and spatial multi-omics as emerging frontiers. Conventional bulk omics obscure cellular heterogeneity within early lesions. Single-cell RNA sequencing and single-cell ATAC-seq simultaneously profile gene expression and chromatin accessibility at individual cell resolution, identifying rare pre-malignant subpopulations and revealing the molecular changes that precede tumor formation. Spatial transcriptomics and spatial proteomics map these molecular profiles to physical locations within tissue sections, preserving the architectural context lost when tissue is dissociated.

Toward personalized screening frameworks. The ultimate vision is a precision prevention model in which each high-risk individual receives a personalized risk assessment based on their unique molecular profile - integrating multi-omics liquid biopsy, imaging, clinical risk factors, and microbiome data - that dynamically updates with each surveillance visit. AI-driven digital twin models of tumor evolution could eventually predict which high-risk nodules will progress and which will remain stable, enabling truly individualized screening intervals and intervention thresholds rather than population-average schedules.

TL;DR: Major barriers including data standardization, demographic gaps, and model interpretability must be addressed, while single-cell spatial omics and federated AI represent the technological frontier for personalized, precision lung cancer screening.
Citation: Open Access, 2026. Available at: PMC12877424.