Lung cancer is the leading cause of cancer-related death. With an estimated 124,730 deaths in the United States in 2025 alone and a 28% five-year survival rate, lung cancer represents one of the greatest unmet needs in oncology. Early-stage detection offers dramatically better outcomes, with 5-year survival reaching 90% for the smallest Stage IA tumors.
More than 1.5 million Americans have lung nodules detected incidentally each year. These nodules are found on CT scans performed for unrelated reasons such as emergency room visits or cardiac imaging. While most nodules are benign, a meaningful subset are malignant and require timely follow-up to achieve curative treatment.
Current clinical management fails most patients with incidental nodules. Approximately 6 out of 10 patients with incidentally detected lung nodules are lost to follow-up. This deficiency in care coordination and risk stratification leads to delayed diagnoses, invasive procedures on benign nodules, and suboptimal outcomes for patients who do have cancer.
This study evaluated whether AI tools could close the follow-up gap. The central question was whether combined AI-based tools for patient identification, risk stratification, and tracking could bring about an earlier stage at diagnosis by ensuring appropriate follow-up for all patients with clinically relevant lung nodules.
The AI tool comprised three integrated modules working in sequence. Patient discovery used automated Natural Language Processing of radiology reports deployed across the entire hospital system to identify patients with suspicious nodules and accelerate specialist referral for high-risk cases. Clinical decision support computed an imaging AI and radiomics-based digital biomarker of malignancy risk directly from CT image pixels. Management and tracking assisted physicians in longitudinally monitoring decisions and reminded care teams when patients missed prescribed follow-up scans.
The study used a retrospective cohort of 65,039 CT radiology reports. All reports from July 2017 to February 2018 at a large academic medical center were processed by the NLP patient discovery module. The study period was chosen to allow a full two-year follow-up window without intersecting with COVID-19-related practice changes.
Inclusion criteria focused on incidentally detected nodules of clinical relevance. Nodules measuring 8 to 30 mm that were not fully calcified were automatically considered clinically relevant. Nodules of 5 to 7 mm were included if the radiology report contained an explicit follow-up recommendation. Patients with a pre-existing cancer diagnosis at the time of the scan were excluded.
A multi-reader multi-case reader study enhanced the analysis. Pulmonologists from outside the academic medical center reviewed 153 CT cases selected from the cohort. Their recommendations were used to model what follow-up decisions would have been made with AI assistance, enabling a paired comparison between standard care and the hypothetical AI intervention.
AI increased guideline-concordant follow-up from 34% to 94%. In current clinical practice, only 471 of 1,393 patients with clinically relevant nodules received follow-up in accordance with established guidelines. With AI, this number would have risen to 1,308 patients, a statistically highly significant improvement confirmed by a bootstrapped McNemar test with p less than 0.0001.
Nearly two-thirds of patients currently receive no nodule-related follow-up. Of the 1,393 patients identified, 881 (63%) received no nodule-related visits within the two-year observation period. Only 41 (3%) received delayed follow-up. This staggering rate of non-follow-up demonstrates the scale of the clinical failure that AI could address.
AI also accelerated the speed of initial follow-up. Time-to-event analysis showed that patients would have been followed up significantly more rapidly with AI tools compared to standard care (log-rank test p less than 0.0001). Smaller nodules were disproportionately more likely to be missed or delayed, a pattern that AI-assisted tracking would specifically correct.
For patients with cancer, AI raised guideline adherence from 56% to 96%. Among the 55 patients estimated to have lung cancer in the cohort, only 31 (56%) were followed according to guidelines in standard care. With AI, this number would have increased to 53 (96%), representing a nearly complete capture of cancer patients who need timely diagnosis and treatment.
Median time to lung cancer diagnosis fell from 129 days to 25 days with AI. In current clinical practice, the mean time to diagnosis was 282 days with a median of 129 days. With AI, the modelled mean was 53 days and the median was 25 days, a statistically significant reduction confirmed by Wilcoxon signed-rank test with p less than 0.001.
AI increased the proportion of cancers diagnosed within 60 days. In standard care, 36% of confirmed cancer patients received histological confirmation within the recommended 60-day window. With AI, 82% would have been diagnosed within that period (p less than 0.01). Earlier histological confirmation enables faster initiation of curative treatment.
Tumor size at diagnosis was modestly reduced by AI. The median maximal axial diameter at diagnosis in current practice was 22 mm. With AI, the modelled median was reduced to 18 mm (p less than 0.01). Smaller tumors at diagnosis are associated with less advanced disease and more options for surgical resection.
Two Stage IV cancers would have been down-staged with AI. Among 22 cancers with known staging, two patients diagnosed at Stage IV under current care would have been diagnosed at a lower stage with AI. Although the small sample limits statistical significance, this finding illustrates the potential for AI to shift late-stage diagnoses to earlier, more treatable stages.
The 60% non-follow-up rate confirms the scale of the unmet need. The finding that approximately 63% of patients with clinically relevant nodules received no follow-up aligns with prior published reports. This underscores the substantial potential benefit of NLP-based AI tools that can automatically flag at-risk patients who would otherwise fall through the gaps in care coordination.
Increased follow-up must be paired with effective risk stratification. Enhanced patient identification will increase the volume of patients referred for specialist evaluation, potentially overwhelming healthcare providers. Risk stratification components of the AI tool that correctly reclassify benign cases from intermediate or high risk to low risk are therefore essential to prevent unnecessary procedures for non-cancerous nodules.
The retrospective design introduces several important limitations. Patients who were never followed could not have their cancer status definitively determined. Some may have been followed at other institutions, died from unrelated causes, or refused treatment for personal or financial reasons. In such cases, AI tools would not improve follow-up, meaning the observed improvements likely represent an upper-bound estimate.
Community settings may benefit even more than academic medical centers. This study was conducted at an academic center with subspecialty-trained thoracic radiology experts. The greatest unmet need exists in regional and community hospitals where generalist radiologists interpret scans and pulmonologists lack lung cancer specialization. AI tools have been shown to reduce variability among readers and may be especially valuable in these lower-resource settings.
The combined AI tool addresses three distinct failure points in nodule management. Patient identification through NLP captures patients whose nodules would otherwise be missed. Patient tracking prevents loss to follow-up by alerting care teams to missed appointments. Patient stratification ensures the most appropriate type and timing of follow-up, avoiding both over- and under-treatment.
Each component of the AI tool serves a distinct and complementary function. NLP-based discovery increases the pool of identified patients. Radiomics-based malignancy risk scoring helps clinicians make more accurate and consistent risk assessments than existing statistical models or unaided clinical judgment. Tracking reminders address the coordination failures that cause patients to fall out of care pathways after initial identification.
Economic considerations support AI implementation. Preliminary cost-effectiveness analyses show a favorable incremental cost-effectiveness ratio when the costs of increased follow-up are weighed against the life years gained through earlier diagnosis. More comprehensive health economic studies are needed to fully quantify these benefits across healthcare systems.
Prospective real-world validation is the essential next step. This retrospective study demonstrates compelling modelled benefits, but the ideal scenario assumed all patients offered follow-up actually received it. A before-and-after prospective study in a real clinical setting is required to confirm these findings and to capture the full stage-shift potential of combined AI tools including those patients currently never identified or followed.