Is Automatic Tumor Segmentation on Whole-Body 18F-FDG PET Images a Clinical Reality?

Journal of Nuclear Medicine 2024 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Why Automated Tumor Segmentation on PET Matters

18F-FDG PET/CT has established itself as an indispensable tool in oncologic imaging, enabling clinicians to detect and monitor tumors by visualizing areas of elevated glucose metabolism. By highlighting regions of increased glucose consumption, this technique allows discrimination of malignant from benign tissue, providing critical information for diagnosis, staging, and evaluation of therapeutic response across a broad range of cancers including lymphoma, lung cancer, and melanoma.

Volumetric metrics: Beyond simple tumor detection, quantifying tumor burden through metrics such as metabolic tumor volume (MTV), disease dissemination index (DDI), and total lesion glycolysis (TLG) has considerably refined the clinical utility of PET imaging. These parameters encapsulate both the volume and metabolic intensity of tumors, serving as indicators of tumor aggressiveness and response to therapy. High MTV and TLG values at baseline are consistently associated with inferior progression-free and overall survival in lymphoma, providing prognostic information that standard clinical indices like the International Prognostic Index (IPI) do not fully capture.

The manual segmentation bottleneck: Despite this prognostic value, manual segmentation of tumor volumes on whole-body PET/CT is labor-intensive and subject to substantial interobserver variability, limiting its feasibility in routine clinical practice. Delineating dozens of hypermetabolic lesions scattered across the body requires significant physician time and produces results that vary depending on the reader, the thresholding method used, and the imaging system. This variability compromises the reproducibility of volumetric biomarkers and makes serial monitoring across treatment cycles unreliable.

This 2024 editorial in the Journal of Nuclear Medicine, authored by Lalith Kumar Shiyam Sundar and Thomas Beyer from the Medical University of Vienna, critically examines whether automated whole-body tumor segmentation driven by artificial intelligence has crossed the threshold from a research curiosity to a clinical reality. The authors survey the academic landscape, review commercial tools, and identify the remaining barriers to routine deployment.

TL;DR: 18F-FDG PET/CT metabolic tumor volume (MTV), total lesion glycolysis (TLG), and disease dissemination index are validated prognostic biomarkers in lymphoma and other cancers, but manual segmentation is too labor-intensive and variable for routine use. This editorial evaluates whether AI-driven automated segmentation is now clinically viable.
Pages 1-2
Academic AI Research for Whole-Body PET Tumor Segmentation

The academic community has invested heavily in leveraging artificial intelligence for automated whole-body PET 18F-FDG tumor segmentation. This transition seeks to overcome the fundamental limitation of traditional thresholding techniques, which segment all hypermetabolic regions above a fixed SUV cutoff without distinguishing pathologic from physiologic tissue uptake. Fixed-threshold approaches indiscriminately capture normal structures such as the brain, heart, liver, and kidneys alongside true tumor lesions, generating false positives that require manual correction.

Deep learning architectures: AI algorithms, particularly those built on the foundational U-Net architecture and its variants, represent the dominant approach in academic research. U-Net's encoder-decoder structure with skip connections was originally designed for biomedical image segmentation and has proven highly effective for detecting lesions in complex 3D imaging volumes. More advanced variants including nnU-Net (a self-configuring U-Net that automatically adapts preprocessing, architecture, and training to any given dataset) and MONAI's Auto3Dseg framework have emerged as state-of-the-art tools offering robust performance across multiple segmentation benchmarks.

Open-source training datasets: Two academic initiatives have been pivotal in advancing the field by providing openly available annotated PET/CT datasets. The AutoPET challenge provides a whole-body FDG-PET/CT dataset with manually annotated tumor lesions drawn from lymphoma, lung cancer, and melanoma patients. The HECKTOR challenge focuses on head and neck cancer, providing PET/CT data with expert lesion annotations for training and evaluation. Both datasets have enabled numerous groups to train, compare, and refine AI segmentation models using standardized benchmarks, accelerating progress that would have been impossible with siloed institutional datasets alone.

Despite the volume of published academic work demonstrating AI-driven segmentation success, the authors note a meaningful gap: there is a notable scarcity of open-source, deployable segmentation solutions. Most published models are not publicly released in usable form, and the available open-source datasets are limited to specific cancer types, restricting generalizability across the full spectrum of oncologic indications encountered in clinical practice.

TL;DR: Academic AI research uses U-Net-based architectures, most notably nnU-Net and MONAI Auto3Dseg, for whole-body PET tumor segmentation. AutoPET (lymphoma, lung cancer, melanoma) and HECKTOR (head and neck cancer) provide the primary open training benchmarks. Despite substantial research activity, open-source deployable tools remain scarce and datasets cover only a narrow cancer type range.
Pages 2-3
What Commercial Vendors Are Offering Today

In response to growing clinical demand for automated lesion quantification, leading commercial vendors have developed fully automated and semiautomated methodologies for whole-body tumor volume segmentation from 18F-FDG PET images. These commercial solutions follow a similar initial step: threshold-based segmentation to highlight all hypermetabolic regions, which necessarily captures both pathologic tumor lesions and physiologic tissue uptake. The key differentiator among vendors lies in how they handle the subsequent task of distinguishing these two tissue types.

Siemens Healthineers Auto ID: Auto ID takes a fully automated approach by applying an AI algorithm to differentiate pathologic from physiologic tissues after the initial threshold-based segmentation step. Rather than requiring the physician to manually identify and remove normal-tissue regions, the AI classifier assigns each segmented region to either the pathologic or physiologic category. This represents a step toward complete automation, reducing the need for manual intervention while still leveraging an initial thresholding step to narrow the candidate lesion set.

Hermes Medical Solutions and MIM Medical: Both companies have adopted a simplified semiautomatic one-click methodology rather than full automation. Hermes offers Single Click Segmentation and MIM Medical provides Lesion ID. In both systems, a threshold-based pre-segmentation is performed first, and the user then manually identifies nonpathologic regions within the pre-segmented result with a single click. The system categorizes regions as either pathologic or nonpathologic based on this user interaction, effectively isolating tumor volumes with minimal manual effort. This approach trades full automation for simplicity and clinical acceptability, giving clinicians direct oversight of the segmentation process rather than delegating the physiologic-vs-pathologic decision entirely to an AI algorithm.

The authors characterize these commercial tools as "smart simplicity": they provide certified solutions that balance efficiency with practicality and offer clinicians straightforward segmented regions for clinical use. However, the authors also note that full automation, where no user interaction is required at all, remains an aspirational goal rather than a current commercial reality for all vendors.

TL;DR: Three commercial platforms offer PET tumor segmentation: Siemens Auto ID (AI-driven fully automated pathologic-vs-physiologic classification), Hermes Single Click Segmentation, and MIM Lesion ID (both semiautomated one-click approaches). All begin with threshold-based pre-segmentation. Full zero-interaction automation is not yet universally achieved across commercial products.
Pages 2-3
Volumetric PET Parameters in Clinical Practice: Promising but Underused

Extensive clinical research has validated the prognostic significance of volumetric parameters extracted from 18F-FDG PET/CT, with metabolic tumor volume and total lesion glycolysis demonstrating robust associations with treatment outcomes across a diverse spectrum of cancers. These metrics provide complementary information to the simpler SUVmax measurement by capturing not just peak metabolic activity but the full spatial extent and total metabolic burden of disease. In lymphoma specifically, MTV and TLG at baseline and after initial treatment cycles are among the strongest independent predictors of progression-free and overall survival, outperforming traditional clinical prognostic indices in many studies.

The clinical adoption gap: Despite this body of evidence, the authors report based on their direct correspondence with nuclear medicine clinicians and vendors that volumetric PET parameters are not being used in routine clinical practice or in the majority of clinical trials. The disconnect between research validation and clinical implementation reflects the manual segmentation barrier: computing MTV and TLG accurately requires delineating every metabolically active lesion across the entire body, a task that is simply too time-consuming for routine clinical workflow even when the prognostic value is well established.

Lymphoma as the primary use case: Among the oncologic indications being tracked for clinical interest in volumetric parameter adoption, lymphoma patient management stands out. There is growing interest from the clinical community in extracting volume-based metabolic parameters specifically for lymphoma, driven in part by the heterogeneous distribution of disease in DLBCL, follicular lymphoma, and other subtypes. In lymphoma patients with extensive, multifocal disease, manual corrections to threshold-based segmentations become particularly tedious, making fully automated solutions especially attractive. This represents both a compelling near-term clinical use case for AI segmentation tools and a driver of commercial investment in the space.

TL;DR: MTV and TLG are research-validated prognostic biomarkers in lymphoma and other cancers but are not in routine clinical use due to the manual segmentation burden. Lymphoma patient management is identified as the primary near-term driver of clinical demand for automated volumetric PET parameter extraction.
Pages 2-3
The Generalizability Problem: Cancer Types, Scanners, and Tracers

The question of whether complete AI automation of whole-body tumor segmentation is achievable encounters significant technical challenges, most centrally the problem of model generalizability. An AI algorithm trained on 18F-FDG PET/CT images from lung cancer patients may not perform equivalently when applied to colorectal cancer, head and neck cancer, or lymphoma. The spatial distribution of disease, the degree of physiologic background uptake in adjacent tissues, the characteristic metabolic intensity of different tumor types, and the morphological appearance of lesions all vary substantially across cancer types. Initial evidence does suggest that algorithms designed for lung cancer segmentation may be adaptable for breast cancer when 18F-FDG PET/CT is used, but these cross-cancer adaptations remain preliminary and have not been systematically validated at the scale needed for clinical deployment.

Scanner and protocol variability: Generalizability challenges extend beyond cancer type to differences in imaging systems and reconstruction protocols across clinical sites. PET images are fundamentally affected by the scanner model, acquisition parameters (particularly injected dose, uptake time, and acquisition duration per bed position), and the reconstruction algorithm used. An AI model trained on data from one vendor's PET/CT system using a specific reconstruction protocol may fail to generalize to images acquired on a different system with different parameters, even within the same cancer type and patient population. This cross-site variability is one of the most practically challenging barriers to multicenter AI deployment.

Tracer diversity: Nuclear medicine uses a broad spectrum of radiotracers beyond 18F-FDG, including PSMA-targeted tracers for prostate cancer, somatostatin analogs for neuroendocrine tumors, and DOTATATE for paragangliomas. Each tracer produces a distinct biodistribution pattern that alters both the signal of interest and the background physiologic uptake that must be suppressed. Developing a separate validated AI model for each tracer imposes a significant economic burden, as clinics face escalating costs with each new model introduction from vendors. The ideal solution would be a unified foundational model capable of handling multiple tracers, but achieving regulatory approval for such a tool is particularly challenging given the specificity of intended uses required in certification processes.

TL;DR: AI segmentation generalizability is limited by cancer type, scanner and reconstruction protocol variability, and tracer diversity. Models trained on lung cancer FDG-PET may show some adaptability to breast cancer but cross-cancer validation is incomplete. Different PET tracers (PSMA, DOTATATE, FDG) each require distinct AI models, increasing economic burden on clinical sites.
Pages 2-3
Dice Scores vs. Clinical Endpoints: A Misaligned Research Priority

A recurring theme in the editorial is the disconnect between how AI segmentation models are currently evaluated in academic research and what actually matters clinically. Current PET-based AI algorithm development places heavy emphasis on optimizing the Dice similarity coefficient (DSC), a geometric overlap metric comparing the AI-generated segmentation against a manually annotated ground truth. On the AutoPET leaderboard, top-performing algorithms report DSC scores of approximately 0.37, reflecting the extreme difficulty of whole-body multi-lesion PET segmentation in a heterogeneous patient population. On HECKTOR, which focuses on the more constrained task of head and neck tumor segmentation, top DSC values reach 0.79, with evaluation criteria that additionally incorporate prediction of clinical outcomes such as overall survival and progression-free survival.

The clinical sufficiency question: The authors argue that optimizing for DSC may not be sufficient for clinical applicability in PET tumor segmentation. A DSC of 0.70 might provide prognostic accuracy for MTV and TLG computation that is comparable to manual segmentation, suggesting that beyond a certain accuracy threshold, further technical improvements in DSC do not translate into meaningful improvements in clinical endpoints. This argument has important implications for research priorities: the field may be expending effort on marginal DSC improvements that do not yield clinically significant gains, while underinvesting in studies that directly evaluate whether AI-derived volumetric parameters improve patient management decisions or outcomes.

Minimum accuracy thresholds: The authors call for a strategic shift in research priorities toward identifying the minimum accuracy threshold that meaningfully enhances clinical endpoints. This would involve prospective studies correlating AI segmentation accuracy, measured not just by DSC but by the downstream accuracy of MTV and TLG calculations, with prognostic stratification performance and treatment response assessment. Such studies would define a clinical sufficiency criterion that algorithm development should target, rather than optimizing indefinitely for geometric overlap metrics that may be poorly correlated with clinical value.

TL;DR: AutoPET top DSC scores reach only ~0.37; HECKTOR top DSC reaches 0.79. The authors argue that a DSC of 0.70 may already provide clinically sufficient prognostic accuracy for MTV/TLG computation, and that further geometric accuracy improvements beyond this threshold may not translate to better clinical outcomes. Research should focus on defining minimum clinical accuracy thresholds rather than chasing DSC leaderboard performance.
Pages 2-3
Dataset Scarcity, Foundational Models, and Regulatory Hurdles

The availability of comprehensive, well-curated PET image datasets remains surprisingly limited given that PET imaging has been in clinical use for decades. In contrast, fields like radiology and pathology have seen rapid AI progress partly driven by the open-sourcing of large annotated image databases. PET imaging has lagged behind, partly because of privacy concerns and the added complexity of managing paired PET/CT data with multisite annotations, and partly because of a lack of coordinated infrastructure and funding for expert lesion labeling at the scale needed for robust AI training. The authors identify creating extensive, publicly available PET image databases as one of the highest-priority needs for the field.

Foundational models and segment anything: Recent advances in computer vision offer a promising direction. Foundation models such as the Segment Anything Model (SAM), trained on massive image datasets to enable flexible segmentation through point prompts, bounding box inputs, or free-form prompts, are already being explored in medical imaging. MedSAM, a medical imaging adaptation of SAM, and commercial applications from companies such as United Imaging Healthcare represent early attempts to apply these general-purpose foundation models to clinical segmentation tasks. The appeal of this approach is that a single large model trained on diverse data could handle multiple cancer types, imaging tracers, and acquisition protocols without requiring separate training runs for each combination.

Regulatory challenges: Achieving regulatory clearance for AI-based segmentation tools adds another layer of complexity. Current certification processes require specific intended-use designations, making it difficult to certify a single unified tool for segmentation across multiple cancer types, tracers, and imaging protocols. Each intended use effectively requires its own validation and approval pathway, which multiplies the regulatory burden. This fragmentation means that even a technically versatile AI tool may need to be approved as separate software versions for each clinical indication, limiting the practical feasibility of the unified foundational model approach that the authors and others in the field envision.

TL;DR: Comprehensive PET image databases remain scarce despite decades of clinical PET use. Foundation models such as SAM and its medical derivative MedSAM offer a path toward versatile multi-cancer segmentation. Regulatory certification requires intended-use specificity, fragmenting approval pathways and making unified multi-indication AI tools difficult to certify under current frameworks.
Page 3
A Collaborative Roadmap for Clinical AI Segmentation

The authors conclude that progress in automated PET tumor segmentation cannot be achieved by any single entity working in isolation. The path forward requires explicit, structured collaboration across academia, industry, and clinical practitioners, each contributing distinct capabilities. Clinicians bring domain expertise in what volumetric parameters are clinically meaningful, what accuracy thresholds are acceptable for different use cases, and how AI tools would realistically integrate into oncology workflows. Without this clinical input, technical development risks optimizing for metrics that do not translate to patient benefit.

Data infrastructure and annotation: Clinicians are also positioned to initiate and contribute to large-scale open-source databases with high-quality annotated data, including standardized metadata on cancer type, disease stage, imaging system, reconstruction protocol, and lesion annotations produced according to consensus guidelines. The importance of standardized annotations cannot be overstated: inconsistent annotation practices produce noisy training labels that limit model accuracy and make cross-dataset comparison impossible. Targeted grant support for professional labeling services is identified as a practical mechanism for accelerating high-quality dataset creation at scale.

Open-source tooling from academia: Academia can contribute through rapid innovation and development of open-source low-click annotation tools such as MedSAM and MONAILabel that reduce the burden of manual annotation for building training datasets. Recent academic work has demonstrated that relatively modest changes, including hyperparameter tuning, minor architecture adjustments, and data augmentation strategies, can produce meaningful improvements in segmentation accuracy, and sharing these findings openly accelerates collective progress. The TMTV-Net model, a fully automated tool specifically designed for total metabolic tumor volume segmentation in lymphoma PET/CT, exemplifies an academic contribution with direct clinical relevance validated in a multi-center generalizability analysis.

Academia-industry partnerships: The MONAI framework, co-created by NVIDIA and King's College London, is highlighted as a model for productive academia-industry collaboration that yields durable software solutions serving both research and clinical needs. The authors advocate for research agreements that go beyond data sharing to include joint development of generic software frameworks, with clearly defined intellectual property rights and monetization strategies established from the outset. Without these structural agreements, academic software frequently suffers from the transient nature and maintenance challenges that arise from a lack of ongoing support incentives, limiting the durability of research outputs even when the underlying science is strong.

TL;DR: The roadmap requires tripartite collaboration: clinicians providing domain expertise and annotated datasets with standardized metadata, academia contributing open-source annotation tools (MedSAM, MONAILabel) and algorithm innovations like TMTV-Net, and industry providing funding and deployment infrastructure. The MONAI framework (NVIDIA and King's College London) is cited as a model for sustainable academia-industry software collaboration.