Enhancing Lymphoma Diagnosis and Follow-Up Using 18F-FDG PET/CT: AI and Radiomics Analysis

Cancers 2024 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-3
Why 18F-FDG PET/CT and AI Are Reshaping Lymphoma Care

Lymphoma encompasses a broad spectrum of immune system malignancies arising from the clonal proliferation of lymphocytes. Non-Hodgkin lymphoma (NHL) accounts for roughly 90% of all lymphoma cases, with Hodgkin lymphoma (HL) comprising the remaining 10%. In 2023, NHL was projected to account for 4.1% of new cancer diagnoses and 3.3% of cancer-related deaths in the United States. The heterogeneous biology of lymphoma, and its ability to mimic post-infectious and inflammatory conditions, makes accurate early diagnosis, risk stratification, and treatment response assessment persistently difficult.

The role of 18F-FDG PET/CT: Positron emission tomography combined with computed tomography using [18F]-fluorodeoxyglucose (18F-FDG PET/CT) is the standard noninvasive three-dimensional imaging modality for lymphoma management. Its indications span initial staging before treatment, re-staging at disease progression, evaluation after therapy, monitoring therapy progress, post-therapy follow-up, and assessing disease transformation. Despite its clinical centrality, 18F-FDG PET/CT has important limitations: variable FDG avidity across lymphoma subtypes, false-negative results in low-grade lymphomas, and false-positive uptake from post-immunochemotherapy inflammation and infection.

Radiomics and AI as the new frontier: Radiomics transforms PET/CT images into high-dimensional quantitative data by extracting hundreds of mathematical descriptors encoding shape, intensity distribution, and texture from segmented tumor volumes. These features capture biological information about tumor heterogeneity, metabolic activity, and spatial organization that lies beyond what visual assessment alone can detect. Artificial intelligence algorithms, including machine learning classifiers and deep learning architectures, can then integrate these radiomic features with clinical and genomic data to build predictive and diagnostic models.

This 2024 review, published in Cancers, provides a comprehensive synthesis of the current literature on the application of AI and radiomics to 18F-FDG PET/CT images in lymphoma. It was conducted by a multidisciplinary team from Shahid Beheshti University, the University of British Columbia, and Geneva University Hospital. The scope covers five clinical domains: bone marrow involvement diagnosis, histologic subtype differentiation, progression and outcome prediction, radiogenomics, and AI-driven segmentation and staging tools.

TL;DR: NHL accounts for 4.1% of new US cancer diagnoses in 2023. 18F-FDG PET/CT is the standard staging and response imaging tool but has variable FDG avidity and false-positive limitations. This 2024 Cancers review synthesizes AI and radiomics applications across five lymphoma management domains, drawing on 128 studies selected from an initial pool of 227 PubMed results spanning 2000-2024.
Pages 3-4
How the Literature Was Selected and Organized

The authors conducted a systematic search of PubMed using Boolean combinations of the terms "lymphoma," "artificial intelligence," "machine learning," "radiomics," "deep learning," and "radiogenomics." The search covered publications from the year 2000 through 20 August 2024. An initial yield of 227 articles was reduced to 128 relevant papers after applying predefined inclusion and exclusion criteria. Exclusion criteria included publications covering a wide range of diseases without lymphoma-specific focus, conference papers, literature reviews, and articles in non-English languages.

Study design distribution: A critical finding from this bibliometric assessment is that up to 95% of the included studies employed a retrospective design. Only one study, Ceriani et al., incorporated a prospective component in its validation phase. This preponderance of retrospective data is flagged by the authors as a central limitation of the field, because retrospective analyses are vulnerable to selection bias and recall bias that may not reflect real-world clinical performance.

Topic distribution: The 128 studies were categorized into five clinical domains according to their primary scientific question. The largest category, comprising roughly 60% of the research, addressed progression and outcome prediction. Histologic subtype differentiation represented approximately 13% of studies, imaging-based biomarker prediction 12%, system-aided diagnosis and staging 7%, and genomics integration the smallest share at about 3%. This distribution reflects the field's maturity, where prognostic modeling with radiomics is furthest along, and radiogenomics remains an early-stage area.

Publication trend: The review documents a significant acceleration in publications combining "Lymphoma," "PET," and "Radiomics" from 2020 onward, coinciding with broader adoption of deep learning tools and the availability of open-source radiomics software such as PyRadiomics, LifeX, and RaCat. No relevant articles were identified on PubMed prior to 2013, confirming this is an emerging field that has grown almost entirely within the past decade. A recognized limitation is the restriction to a single database (PubMed), which may have excluded relevant studies indexed in EMBASE, Scopus, or Cochrane.

TL;DR: 227 PubMed articles screened, 128 selected. Up to 95% are retrospective. Outcome prediction dominates the literature (60% of studies), followed by histologic differentiation (13%) and biomarker prediction (12%). Publication volume surged after 2020. Single-database search is acknowledged as a limitation.
Pages 5-7
Radiomics for Detecting Bone Marrow Infiltration Without Biopsy

Bone marrow infiltration (BMI) is clinically critical in lymphoma because it directly affects staging, prognosis, and therapeutic management. The gold standard for BMI assessment has historically been bone marrow biopsy (BMB), an invasive procedure subject to sampling error and patient discomfort. 18F-FDG PET/CT is increasingly recognized as a staging tool with potential to serve as a non-invasive surrogate for BMB, and radiomic feature extraction from PET/CT images can further enhance this capability. The importance of BMI varies considerably across subtypes: B-cell lymphoblastic lymphoma 40-60%, small lymphocytic lymphoma 85%, mantle cell lymphoma (MCL) 50-80%, follicular lymphoma (FL) 40-70%, aggressive NK cell leukemia greater than 95%, and lymphocyte-depleted subtype greater than 50%.

Key radiomic studies: In 66 FL patients, Faudemer et al. demonstrated that FDG-PET/CT radiomics achieved AUC 0.822, sensitivity 70.0%, and specificity 83.3% for BMI diagnosis, with predictive features including variance and correlation from the grey-level co-occurrence matrix (GLCM) and busyness from the neighboring grey-level dependence matrix (NGLDM). Critically, there was a significant difference between BMB and visual PET assessment (p = 0.010), but no significant difference between BMB and the radiomic prediction score, suggesting radiomics may close the gap between non-invasive imaging and biopsy-level accuracy.

DLBCL and skeletal texture: In 82 DLBCL patients, Aide et al. identified that the radiomic feature SkewnessH, a first-order metric capturing the asymmetry of the intensity histogram, demonstrated sensitivity of 81.8% and specificity of 81.7% for detecting BMI on baseline FDG PET/CT scans. This feature also served as an independent predictor of progression-free survival (PFS). In MCL patients, FDG-PET texture features significantly enhanced SUV-based BMI prediction (AUC 0.73 vs. 0.66), and combining radiomic features with laboratory data further improved performance to AUC 0.81.

AI-based quantification: Sadik et al. proposed an innovative AI-based approach to detect diffuse bone marrow uptake (BMU) in HL patients, achieving an 81% concordance rate between AI and physician assessments. In a follow-up study, the same group demonstrated that AI assistance significantly enhanced interobserver agreement among physicians from different hospitals, with mean kappa improving from 0.51 without AI to 0.61 with AI guidance. This reduction in interobserver variability is clinically meaningful because inconsistent BMI assessments directly affect staging and treatment decisions.

TL;DR: Radiomic features from PET/CT reach AUC 0.73-0.822 for BMI detection across FL, DLBCL, and MCL subtypes, with sensitivity up to 81.8% and specificity up to 83.3%. Adding laboratory data to radiomics boosts MCL BMI prediction to AUC 0.81. AI-assisted BMU detection achieves 81% physician concordance and improves interobserver kappa from 0.51 to 0.61 in HL patients.
Pages 6-8
Using PET/CT Radiomics and Deep Learning to Differentiate Lymphoma Subtypes

The current gold standard for lymphoma diagnosis is pathological examination of biopsied tissue, but this approach carries known pitfalls: insufficient tissue acquisition, subjectivity in interpretation, errors in molecular genetic testing, and the inability of standard biopsies to capture tumor heterogeneity in relapsed or refractory cases. Radiomic and machine learning approaches applied to PET/CT images offer a complementary, non-invasive route to subtype characterization. Different subtypes of bulky mediastinal lymphomas, including classical HL (CHL), gray zone lymphoma, and primary mediastinal B-cell lymphoma (PMBCL), exhibit distinct FDG-PET characteristics and metabolic heterogeneity that ML-guided radiomics can help differentiate.

FL versus DLBCL: In a retrospective study of 120 patients by de Jesus et al., a radiomics-driven ML model achieved AUC 0.86 and accuracy 80% for discriminating FL from DLBCL, outperforming SUVmax-based models (AUC 0.79, accuracy 70%). This improvement is meaningful because FL and DLBCL represent fundamentally different clinical entities requiring distinct management strategies, yet they can present with similar imaging phenotypes at the organ level.

HL versus DLBCL: Lovinfosse et al. studied 251 patients and reported that the best ML models for differentiating DLBCL from HL achieved AUC 0.95 in a lesion-based approach using tumor-to-liver ratio (TLR) radiomics, and AUC 0.86 in a patient-based approach combining original radiomics and age. For distinguishing DLBCL from mucosa-associated lymphoid tissue (MALT) lymphomas, total lesion glycolysis (TLG) achieved AUC 0.906, sensitivity 0.900, and specificity 0.780.

Deep learning-assisted lymph node classification: Yang et al. developed deep learning-based computer-aided diagnosis (DL-CAD) systems using PET/CT images to address the clinical challenge of differentiating lymph node metastasis from lymphoma involvement in enlarged cervical lymph nodes, achieving AUC 0.901, accuracy 86.96%, sensitivity 76.09%, and specificity 94.20%. Combining handcrafted radiomic features with DL-based features improved diagnostic performance beyond either approach alone. A recent PST-Radiomics approach captured intra-tumor metabolic heterogeneity (ITMH) from PET/CT using a specialized neural network, outperforming traditional subtype classification methods by analyzing tumor segments at multiple spatial scales.

TL;DR: Radiomic ML models discriminate FL from DLBCL at AUC 0.86 vs. 0.79 for SUVmax alone. HL vs. DLBCL differentiation reaches AUC 0.95 with TLR-based lesion radiomics. DL-CAD for lymph node vs. metastasis classification achieves AUC 0.901. DLBCL vs. MALT lymphoma: TLG reaches AUC 0.906, sensitivity 0.900.
Pages 8-12
Radiomics and Machine Learning for Survival and Treatment Response Prediction

Diffuse aggressive NHL achieves complete remission in 60-80% of cases, but 20-40% of patients experience relapses. HL responds well to chemotherapy in the majority but 5-10% have refractory disease initially and 10-30% relapse. Despite clinical prognostic tools like the International Prognostic Index (IPI), accurate prognosis prediction remains challenging due to significant tumor heterogeneity within each lymphoma subtype. The IPI relies on only five binary clinical variables (age, stage, LDH level, performance status, number of extranodal sites) and provides coarse risk stratification that cannot account for the imaging and molecular diversity within each risk group.

DLBCL outcome prediction: A combined clinical and PET/CT radiomics model demonstrated AUC 0.75 for predicting 2-year event-free survival (2-EFS) under R-CHOP treatment, compared to AUC 0.67 for metabolic tumor volume (MTV) alone. Eertink et al. showed that adding baseline radiomics to clinical predictors improved model performance to AUC 0.79, a 15% increase in positive predictive value compared to IPI alone (AUC 0.68), and identified twice the proportion of high-risk patients who progressed within 2 years (44% vs. 28%). Jiang et al. developed a PET radiomics signature (Rad-Sig) that achieved C-index 0.801 for PFS and 0.807 for OS in DLBCL patients, with external validation confirming C-indices of 0.758 and 0.794, respectively.

Deep learning for treatment failure prediction: A multimodal deep learning model predicting treatment failure in DLBCL from PET/CT imaging achieved 91.22% accuracy and AUC 0.926 in the primary dataset, and maintained strong performance (88.64% accuracy, AUC 0.925) in the external dataset. A CNN model applied to maximum intensity projection (MIP) images from baseline 18F-FDG PET improved 2-year time-to-progression (TTP) prediction compared to IPI (AUC 0.74 vs. 0.68). Using deep learning and rule-based reasoning on pre-treatment CT and PET images, one study surpassed traditional prognostic methods with AUC 0.93 in a lesion-based approach.

Hodgkin lymphoma and other subtypes: For classical HL (cHL) patients treated with salvage chemotherapy and autologous stem-cell transplant, a predictive ML model using radiomics and clinical data from 113 patients achieved AUC 0.810 in the training cohort and 0.750 in the validation cohort. A model using first- and second-order radiomic features in early-stage HL predicted outcome with AUC 95.2%, pointing toward personalized risk-adapted management. For CAR-T cell therapy in relapsed or refractory DLBCL, a combined radiomics and clinical model outperformed the clinical model alone (AUC 0.776 vs. 0.712 for PFS and 0.828 vs. 0.728 for OS in training, with validation AUCs of 0.886 vs. 0.635 for PFS). A multi-center study on 240 DLBCL patients using stacking ensemble learning with PET radiomics demonstrated PFS AUC 0.771 and OS AUC 0.725, offering better risk stratification than clinical factors alone.

TL;DR: IPI alone achieves AUC 0.68 for DLBCL 2-year EFS; combined radiomics plus clinical models reach AUC 0.79, a 15% PPV gain. Deep learning treatment failure models hit AUC 0.926 with 91.22% accuracy. CAR-T outcome prediction improves from AUC 0.712 (clinical alone) to 0.886 (combined) in validation. HL ML model: AUC 0.810 training, 0.750 validation. C-index for Rad-Sig: 0.801 PFS, 0.807 OS with external validation at 0.758 and 0.794.
Pages 13-14
Integrating Genomics, Circulating Tumor DNA, and PET/CT Radiomics

Radiogenomics, the integration of radiomic features with genomic data derived from high-throughput sequencing or gene expression profiling, represents an emerging frontier in precision oncology. It aims to decode the molecular mechanisms underlying the imaging phenotype of tumors, potentially revealing non-invasive surrogates for genetic alterations that currently require invasive biopsy. Despite the conceptual appeal of this approach, published research in the context of lymphoma remains limited, representing only about 3% of the reviewed studies.

Metabolic gene signatures in DLBCL: Mazzara et al. identified a 6-gene metabolic signature (T-GEP) and produced a predictive Rad-Sig by combining FDG-PET radiomics and T-GEP data. This Rad-Sig was significantly correlated with the metabolic gene expression profiling-based signature (r = 0.43, p = 0.0027) and associated with PFS (p = 0.028). This demonstrates that specific radiomic features can serve as surrogate imaging biomarkers for underlying tumor metabolic gene expression patterns, potentially allowing non-invasive assessment of cancer metabolism.

Genomic predictors of treatment response: Ferrer-Lores et al. combined imaging characteristics, clinical factors, and genomic data to predict treatment response in DLBCL patients, achieving a combined model AUC of 0.904 with 90% accuracy. BCL6 amplification emerged as a highly predictive genetic marker (p = 0.018), while GLSZM GrayLevelVariance (p = 0.048), sphericity (p = 0.027), and GLCM correlation (p = 0.05) were radiomic predictors of response. In 24 B-cell lymphoma patients, combining 18F-FDG PET/CT radiomics with genomic factors, including MYC and BCL2 double-expressor (DE) status, successfully stratified patients into three risk groups with 3-year PFS of 85.7%, 63.6%, and 0%, respectively.

Liquid biopsy and ctDNA: Circulating tumor DNA (ctDNA), measured through liquid biopsy, provides genotypic insights and evaluates treatment effectiveness by detecting minimal residual disease (MRD). Quantitative, mutational, and fragmentation features in ctDNA can dynamically predict treatment response and survival. Combining LiqBio-MRD with PET/CT results in strong diagnostic accuracy for identifying FL patients with rapid progression, underscoring the value of liquid biopsy in early high-risk patient identification. The authors identify a significant research gap: very few studies have explicitly combined ctDNA or liquid biopsy data with PET/CT radiomics features, despite both being minimally invasive and enabling repeated longitudinal assessments during treatment.

TL;DR: FDG-PET Rad-Sig correlates with metabolic gene expression signature at r = 0.43 (p = 0.0027) and predicts PFS (p = 0.028). Combined imaging-clinical-genomic model reaches AUC 0.904 and 90% accuracy; BCL6 amplification is a top genetic predictor. PET plus genomic risk stratification separates 3-year PFS into 85.7%, 63.6%, and 0% groups. Integrating ctDNA with PET radiomics is identified as a major unexplored research opportunity.
Pages 22-27
Automated Segmentation, TMTV Calculation, and AI-Assisted Staging

Total metabolic tumor volume (TMTV) and total lesion glycolysis (TLG) are among the most validated prognostic biomarkers in lymphoma, but their manual calculation on whole-body PET/CT scans is time-consuming, subject to interobserver variability, and operationally infeasible for routine clinical use. In DLBCL patients, high baseline TMTV is consistently associated with inferior PFS and OS. CNN-based algorithms have been developed to automate TMTV measurement, with results showing high correlation between automated and manually derived TMTV values across multiple studies.

Automated MTV and TMTV: Capobianco et al. applied a CNN-based PARS system to 301 DLBCL patients, achieving a correlation between automated TMTV (TMTV-PARS) and reference TMTV (rho = 0.76, p less than 0.001), with hazard ratios for PFS of 2.3 and OS of 2.8 using automated TMTV, compared to 2.6 and 3.7 for the manual reference. Kuker et al. demonstrated nearly perfect correlation between automated and manual TMTV measurements (Pearson r = 0.9814-0.9818, ICC = 0.98 for both readers). Karimdjee et al. showed fully automated methods achieve ICC for TMTV of 0.99 and for TLG of 1.0, with processing time reduced to just 20 seconds for the fastest method versus 326 seconds for semi-automated workflows.

Deep learning segmentation: A multi-resolution 3D U-Net model for TMTV segmentation achieved an average Dice score of 0.68 on internal data and 0.66 on multi-center external data, with a correlation of 0.89 to ground truth TMTV. Semi-automated and deep learning-based segmentation approaches have achieved Dice coefficients of 0.8796 to 0.907 in separate studies. For ENKT lymphoma, a computer-aided diagnosis (CAD) system achieved Dice coefficient 0.7115 and sensitivity 0.7472. The Self-Adaptive Configuration (SAC) Bayesian method outperformed other segmentation algorithms with intra-observer Dice coefficient of 0.87 and inter-observer Dice coefficient of 0.94 on 18F-FDG PET/CT lymphoma images.

Staging assistance and Deauville scoring: The Deauville 5-point scale is the standard for assessing treatment response on PET/CT but suffers from interobserver variability, particularly for scores 3 and 4. Sadik et al. introduced an AI-driven method to quantify reference levels in both the liver and mediastinal blood pool, with the validation group showing a mean difference of only 0.02 between AI-based and radiologist measurements, demonstrating near-perfect agreement. A deep learning model effectively differentiated 18F-FDG PET/CT scans with and without hypermetabolic tumor sites (AUC 0.953, accuracy 0.907, sensitivity 0.874, specificity 0.949 in external testing), suggesting its potential as a second-reader or rule-out tool. ChatGPT was also evaluated for answering patient questions about PET/CT scans, providing "appropriate" responses for 92% and "useful" answers for 96% of questions, though 16% of responses were inconsistent.

TL;DR: Automated TMTV via CNN achieves ICC 0.98-0.99 versus manual methods, reducing processing time from 326s to 20s. 3D U-Net segmentation Dice: 0.66-0.68 externally, 0.89 TMTV correlation. SAC Bayesian segmentation achieves intra-observer Dice 0.87, inter-observer 0.94. AI-quantified Deauville reference levels differ from radiologist by only 0.02. Deep learning tumor detection: AUC 0.953, specificity 0.949 in external testing.
Pages 29-31
Barriers to Clinical Translation and the Path Forward

Retrospective design and small sample sizes: The most critical limitation across the reviewed literature is that up to 95% of studies are retrospective. Retrospective designs introduce selection bias and cannot prove that AI tools improve clinical outcomes. Most individual studies are conducted at single centers with datasets ranging from 24 to a few hundred patients, which is insufficient to train and validate generalizable deep learning models for rare lymphoma subtypes. For instance, the ENKTL studies noted that lower radiomic model performance in validation compared to training was likely attributable to the small dataset size and retrospective study design.

External validation gaps and feature instability: Eertink et al. demonstrated that different segmentation techniques produce differences in radiomic feature values and alter which features are selected for predictive models in DLBCL. This feature instability is a fundamental challenge for reproducibility: radiomic features computed with different software (LifeX, PyRadiomics, RaCat, InterView FUSION) or different segmentation parameters can yield substantially different values even for the same patient scan. Standardization of feature extraction pipelines is a prerequisite for multi-institutional comparison and regulatory submission. Performance drops of 5-15% between internal cross-validation and external validation are consistently reported across studies.

Interpretability and clinical workflow integration: Deep learning models operating on raw PET/CT voxel data function largely as black boxes, and Grad-CAM or attention map visualization techniques provide only partial transparency into their decision-making. Clinicians require outputs that align with known biological and staging criteria. AI tools must also integrate seamlessly into existing nuclear medicine workflow software and reporting systems, a non-trivial engineering challenge. Regulatory pathways (FDA clearance or CE marking), liability frameworks, and reimbursement models for AI-assisted interpretation are still immature in most healthcare systems.

Future directions: The authors highlight several specific research priorities. First, federated learning offers a practical solution for training models across institutions without sharing sensitive patient data, enabling rare-subtype training sets that no single center could assemble alone. Second, Vision Transformers (ViT) and attention-based CNNs can enhance quantitative analysis of PET/CT images with better long-range spatial feature capture. Third, a hybrid AI-human collaboration model could combine AI automation with clinical oversight, reducing false-positive rates while maintaining throughput. Fourth, the integration of radiomic features with ctDNA liquid biopsy data represents an unexplored high-potential avenue for non-invasive lymphoma monitoring. Prospective randomized clinical trials comparing AI-assisted versus standard-of-care lymphoma management are urgently needed to generate the evidence required for routine clinical adoption.

TL;DR: 95% of studies are retrospective; external validation drops of 5-15% are routine. Feature instability across segmentation methods and software platforms undermines reproducibility. Key future priorities: federated learning for multi-institutional rare-subtype training, Vision Transformers for PET image analysis, hybrid AI-human collaboration models, ctDNA plus radiomics integration, and prospective randomized trials to establish clinical benefit.