Assessments of Lung Nodules by an Artificial Intelligence Chatbot Using Longitudinal CT Images

Cell Rep Med 2025 AI 5 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
GPT-4o as a Radiologist: Evaluating Lung Nodules from CT Video

The Problem with Current Nodule Assessment Lung nodule evaluation requires expert radiologists to assess size, growth rate, malignancy risk, and morphology from CT scans over time. This is resource-intensive, prone to inter-observer variability, and inaccessible in many settings. AI solutions that can match radiologist-level performance could transform screening programs.

Using a Vision-Language Model (VLM) This study evaluated GPT-4o, a multimodal large language model with vision capabilities, for assessing lung nodules from longitudinal CT images. Unlike specialized deep learning models, GPT-4o is a general-purpose AI chatbot that receives image inputs and produces structured text outputs.

CT to Video Conversion A novel input strategy was used: consecutive CT slices through the nodule were compiled into a short video at 20 frames per second. This allowed GPT-4o to perceive the 3D structure of the nodule across slices, simulating how radiologists mentally reconstruct 3D anatomy from 2D slices.

Study Scope The study evaluated 647 patients across four independent datasets (C1, C2, LLCS, and NLST), testing GPT-4o on malignancy classification, nodule size measurement, and feature detection tasks. This broad validation across multiple real-world cohorts provides robust evidence for or against clinical applicability.

TL;DR: This study tested GPT-4o - a general-purpose AI chatbot - on lung nodule evaluation by converting CT slices into short videos, evaluating 647 patients across four independent datasets.
Pages 3-4
Dataset Design and Video Input Strategy

Four Independent Datasets The study used four cohorts: C1 (institutional dataset with confirmed malignancy labels), C2 (a second institutional dataset), LLCS (a lung cancer screening cohort), and NLST (the National Lung Screening Trial - the landmark US screening study). Using NLST adds particular significance as it is the most widely validated lung screening dataset.

Video Generation Protocol For each nodule, CT DICOM files were processed to extract slices spanning the nodule extent. These were assembled into MP4 video files at 20 fps and paired with baseline and follow-up timepoints, giving GPT-4o a temporal view of nodule evolution - a capability crucial for detecting growth.

Prompt Engineering GPT-4o was queried with structured prompts asking for: (1) malignancy probability (0-100%), (2) nodule diameter measurement in millimeters, (3) presence/absence of specific features (spiculation, lobulation, ground-glass opacity, calcification), and (4) Likert-scale confidence ratings. Prompts were standardized across all cases.

Evaluation Metrics Performance was measured using AUC for malignancy classification, intraclass correlation coefficient (ICC) for size measurement agreement with radiologists, and Likert scores for feature detection quality. These metrics allowed comparison against established radiologist performance benchmarks.

TL;DR: CT slices were converted to 20fps video clips and input to GPT-4o with structured prompts asking for malignancy classification, size measurement, and feature detection across four datasets including NLST.
Pages 5-7
GPT-4o Matches Expert Radiologists on Key Tasks

Malignancy Classification Accuracy GPT-4o achieved an accuracy of 0.88 for malignancy classification across datasets. This is competitive with subspecialty-trained radiologists and significantly above what general radiologists achieve without specialized lung cancer training, demonstrating meaningful clinical-grade performance.

Nodule Size Measurement For nodule diameter measurement, GPT-4o achieved an intraclass correlation coefficient (ICC) of 0.91 compared to radiologist measurements. An ICC above 0.90 is considered excellent agreement, meaning GPT-4o can reliably quantify nodule size - a critical metric for Lung-RADS categorization and growth assessment.

Feature Detection Quality For identifying nodule characteristics (spiculation, lobulation, ground-glass opacity, calcification), GPT-4o achieved a mean Likert score of 4.17 out of 5, indicating strong agreement with expert radiologist assessments. This shows the model can describe not just whether a nodule is suspicious but characterize what makes it suspicious.

Longitudinal Growth Detection GPT-4o demonstrated particular strength in detecting interval growth between baseline and follow-up scans - the key clinical question in nodule surveillance programs. Correctly identifying growing nodules is the most critical task in lung cancer screening, as growth indicates malignant potential.

TL;DR: GPT-4o achieved 88% malignancy classification accuracy, ICC 0.91 for size measurement, and Likert 4.17/5 for feature detection - all competitive with subspecialty radiologist performance.
Pages 8-9
Implications for Lung Cancer Screening Programs

Democratizing Expert Radiology Subspecialty thoracic radiology expertise is concentrated in academic centers. If GPT-4o or similar VLMs can reliably assess lung nodules, screening programs at community hospitals or in low-resource settings could provide expert-equivalent nodule evaluation without requiring specialized radiologist review for every case.

Reducing Radiologist Burden Lung cancer screening programs generate large volumes of CT scans requiring nodule tracking across multiple timepoints. AI-assisted triage - flagging high-confidence benign nodules and prioritizing suspicious ones for expert review - could dramatically reduce workload and turnaround time.

General-Purpose AI vs. Specialized Models A key finding is that a general-purpose commercial AI chatbot (GPT-4o) performs comparably to specialized deep learning models trained specifically on lung CT data. This suggests that large multimodal foundation models may be approaching the capability to serve as generalist medical imaging tools.

Limitations in Current Clinical Practice GPT-4o cannot directly access DICOM files, requires the video conversion preprocessing step, and lacks integration with PACS systems. Regulatory approval and liability frameworks for AI-assisted diagnosis do not yet cover general-purpose LLMs, creating practical barriers to immediate clinical deployment.

TL;DR: GPT-4o's performance suggests general-purpose multimodal AI could democratize expert lung nodule assessment for screening programs, though PACS integration and regulatory pathways remain barriers.
Pages 10-11
Current Limitations and the Path Forward

Prompt Sensitivity GPT-4o's outputs depend heavily on how questions are framed. Different prompt wording can yield different responses for the same image, raising concerns about reproducibility and standardization. Future deployment would require locked, validated prompt templates.

Video Compression Artifacts Converting DICOM slices to video introduces compression artifacts that may degrade fine spatial details. Developing better native multimodal CT interfaces - or training models directly on DICOM data - would improve fidelity and potentially further boost performance.

Lack of Uncertainty Quantification GPT-4o does not provide calibrated confidence intervals or uncertainty estimates for its predictions. Clinical decision-making requires knowing not just what a model predicts but how confident it is and when predictions should be flagged for human review.

Future Directions Next steps include prospective clinical trials comparing GPT-4o against radiologists in real screening workflows, development of standardized CT-to-LLM input pipelines, evaluation of newer GPT versions or open-source multimodal models, and regulatory studies to define appropriate use cases for VLMs in radiology.

TL;DR: Prompt sensitivity, video compression, and lack of calibrated uncertainty are key limitations; prospective clinical trials and standardized input pipelines are needed before clinical deployment.
Citation: Open Access, 2025. Available at: PMC11970393.