Two-Year Event-Free Survival Prediction in DLBCL Patients Based on In Vivo Radiomics

Frontiers in Oncology 2022 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Why Predicting 2-Year Survival in DLBCL Matters

Diffuse large B-cell lymphoma (DLBCL) is the most common histologic subtype of Non-Hodgkin lymphoma (NHL) worldwide, comprising roughly 30-40% of all NHL diagnoses each year. NHL itself accounts for nearly 3% of cancer diagnoses and deaths globally, making it the most common hematological malignancy. Within DLBCL, clinical outcomes are highly heterogeneous: while standard treatment with the R-CHOP-21 regimen (rituximab, cyclophosphamide, doxorubicin, vincristine, and prednisone) achieves durable remission in approximately 50-60% of patients, around 25-30% are resistant to first-line chemo-immunotherapy. For these patients, salvage therapy including high-dose chemotherapy and autologous hematopoietic stem cell transplantation (ASCT) is required, and those who relapse after ASCT have very poor prospects.

The clinical need for early risk identification: Because the majority of refractory or relapsed DLBCL patients will ultimately die from their disease, identifying this high-risk group before treatment begins is a critical clinical priority. Early identification enables clinicians to consider intensified first-line regimens, clinical trial enrollment, or more aggressive monitoring protocols tailored to individual patients. The initial workup for suspected lymphoma involves comprehensive laboratory panels (LDH, beta-2-microglobulin, EBV and hepatitis serology, complete blood counts), excisional lymph node biopsy for definitive subtype diagnosis, and staging by functional imaging.

The role of PET/CT and event-free survival as an endpoint: The preferred imaging modality for DLBCL staging and treatment monitoring is [18F]FDG PET/CT, which captures metabolically active tumor tissue throughout the body. Two-year event-free survival (EFS) - defined as the time from registration to disease relapse, progression, or lymphoma-related death - has been established as a clinically meaningful and robust endpoint for DLBCL. Patients who achieve 2-year EFS have been shown to have subsequent survival outcomes comparable to an age-matched general population, making it a practical prognostic milestone.

This 2022 study from a Hungarian dual-center cohort investigates whether radiomic features extracted from baseline pretreatment [18F]FDG PET/CT scans, combined with clinical parameters, can train an automated machine learning model capable of predicting which patients will and will not achieve 2-year EFS. The study uses one center's data as a training set and the other center's data as an independent external validation set, providing a more rigorous test of generalizability than single-center cross-validation alone.

TL;DR: DLBCL is the most common NHL subtype (30-40% of cases), with R-CHOP curing 50-60% of patients but 25-30% being refractory. This study builds an automated machine learning model from baseline [18F]FDG PET/CT radiomic features to predict 2-year EFS, using a dual-center design (Center 1 = training, Center 2 = independent validation) across 85 patients.
Pages 2-3
Dual-Center Retrospective Cohort: 85 DLBCL Patients

The study enrolled 85 patients diagnosed with DLBCL whose baseline pretreatment [18F]FDG PET/CT scans were performed between January 2014 and December 2019 across two clinical centers in Hungary. Center 1 (University of Pecs, Department of Medical Imaging) contributed 41 patients, and Center 2 (University of Kaposvar) contributed 44 patients. The cohort had a median age of 59 years (range: 23-81), with 48.2% of patients older than 60. Gender distribution was 40 males (47%) and 45 females (53%). Patients with incomplete medical records or those who received non-standard treatments were excluded from analysis.

Treatment uniformity: All patients received standard R-CHOP-21 for at least 4 full cycles, ensuring that the training and validation cohorts were treated with the same protocol and making outcome comparisons valid across centers. Patients were classified into germinal center B-cell-like (GCB) or activated B-cell (non-GCB) subtypes using the Hans algorithm, a widely used immunohistochemical surrogate for gene expression-based cell-of-origin (COO) classification. COO data were available for 82 of the 85 patients: 29 (37.6%) were GCB and 53 (62.4%) were non-GCB.

Clinical staging and prognostic indices: Disease stage was assessed using the modified Ann Arbor and Lugano classification systems. The Revised International Prognostic Index (R-IPI), a key clinical risk tool for DLBCL, was calculated for all patients before therapy initiation. R-IPI scores in this cohort were distributed as follows: score 0 in 8 patients, score 1 in 15, score 2 in 23, score 3 in 27, and score 4 in 12 patients. The Eastern Cooperative Oncology Group (ECOG) performance status was greater than 2 in 27 patients (31.8%), with 2 cases having unknown ECOG status.

Outcome distribution: Two-year EFS was defined as the reference standard. Patients were divided into Group 0 (no events during 2-year follow-up, i.e., sustained remission) and Group 1 (primary refractory disease or relapse within 24 months). Of the 85 patients, 55 achieved complete metabolic remission at the end of induction therapy. During follow-up, 14 patients had primary refractory disease, 14 relapsed within 12 months, and 2 relapsed between 12 and 24 months, yielding 30 patients total in the event group. Chi-square analysis confirmed statistically significant associations between group membership and stage (p = 0.017), R-IPI (p = 0.015), and COO (p = 0.018), with no significant sex difference (p = 0.611).

TL;DR: 85 DLBCL patients (median age 59, 53% female), all treated with R-CHOP-21. Center 1 (n=41) served as training set, Center 2 (n=44) as independent test set. 55 patients achieved complete remission; 30 had events within 24 months. Stage (p=0.017), R-IPI (p=0.015), and COO (p=0.018) were significantly associated with 2-year EFS outcome.
Pages 3-4
PET/CT Acquisition, Tumor Delineation, and Feature Extraction

The two clinical centers used different PET/CT scanners, a design choice that adds real-world variability but also tests cross-scanner generalizability. Center 1 used a Mediso AnyScan 16 PET/CT scanner, with PET images reconstructed using the Tera-Tomo 3D algorithm in a 167x167x234 matrix (isotropic voxel size of 4 mm). Center 2 used a Siemens Biograph Truepoint 64 PET/CT scanner, with images reconstructed using the 2D OSEM algorithm (3 iterations, 8 subsets, 5 mm Gaussian filter) in a 168x168 matrix. Both centers used 120 kVp low-dose CT for attenuation correction, and both administered 3-4 MBq/kg of [18F]FDG intravenously after a minimum 6-hour fast, with blood glucose confirmed below 8 mmol/L.

Lesion delineation and segmentation: Lymphoma lesions were detected and delineated using InterView FUSION version 3.10 clinical evaluation software (Mediso Medical Imaging Systems, Budapest). A semi-automated algorithm using the average SUVmax of the liver (range 3.5-5.5) as a reference threshold was used to minimize patient-specific radiotracer distribution differences. Three randomly placed volumes of interest (VOIs) in unaffected liver regions were averaged to establish the threshold. Non-affected regions with physiological activity (bladder, kidneys, brain) and incidental FDG uptake not related to lymphoma (such as bowel uptake from metformin intake) were manually excluded from the delineation.

Extracted features: From the delineated lesions, conventional PET parameters were automatically calculated across all delineated foci: total lesion glycolysis (TLG), total metabolic tumor volume (TMTV), and SUVmax. SUVpeak values were segmented from the VOI with the highest activity. For radiomics, the largest VOI per patient was selected for deeper feature extraction. The extracted features followed the Imaging Biomarker Standardization Initiative (IBSI) guidelines and covered six feature families: intensity, histogram, morphological, neighborhood gray-tone difference matrix (NGTDM), gray-level co-occurrence matrix (GLCM), gray-level run length matrix (GLRLM), and gray-level size zone matrix (GLSZM). Restricting radiomics analysis to the largest lesion, rather than all lesions, is a deliberate methodological choice consistent with prior DLBCL radiomics studies.

IBSI compliance: Adherence to IBSI standards is a notable methodological strength of this study. The Imaging Biomarker Standardization Initiative provides a consensus framework for radiomic feature definitions and computation, addressing the reproducibility crisis in radiomics where nominally identical features computed with different software packages yield different values. IBSI-compliant reporting makes this study's features potentially replicable and comparable to other IBSI-adherent work, a prerequisite for eventual clinical translation.

TL;DR: Two different PET/CT scanners were used (Mediso AnyScan vs. Siemens Biograph), adding cross-scanner generalizability. Liver SUVmax (3.5-5.5) served as the segmentation threshold. Features extracted per IBSI standards included NGTDM, GLCM, GLRLM, GLSZM families plus conventional metrics (TLG, TMTV, SUVmax). Radiomics analysis was restricted to the largest lesion per patient.
Pages 4-6
Automated Machine Learning: Feature Ranking and Model Building

The Center 1 dataset (41 patients) was used as the training set, selected because it had more balanced subgroups between the remission and progression categories compared to Center 2. The training pipeline used the Dedicaid AutoML Research Package (Dedicaid GmbH, Vienna, Austria), a specialized automated machine learning platform for biomedical prediction tasks. The AutoML pipeline performs multiple preprocessing steps automatically: redundancy reduction to remove highly correlated features, class imbalance reduction to handle the unequal numbers of remission and progression patients, and feature engineering, ranking, and selection to identify the most prognostically informative variables.

100-fold Monte Carlo cross-validation: The Center 1 data were split into 100 folds via random subsampling (Monte Carlo cross-validation). Mixed ensemble learning was applied within each fold to generate a prediction of 2-year EFS. This approach differs from standard k-fold cross-validation: in Monte Carlo cross-validation, the dataset is randomly partitioned into training and validation subsets many times, and each partition is independent. Using 100 folds provides a highly stable estimate of model performance by averaging across many random data splits, reducing the sensitivity to any single fortunate or unfortunate partition.

Feature importance ranking: The final feature ranking was generated by averaging feature importances across all 100 folds and normalizing them to sum to 1.0. Features with importance greater than half of the highest-ranked feature were designated as high-ranking and analyzed as candidate imaging biomarkers. This threshold-based filtering ensures that only features with consistent importance across the 100 Monte Carlo folds are retained, reducing the influence of features that happen to be important in only a subset of data partitions.

Top 5 predictive features identified: The automated analysis identified five features as the most important for 2-year EFS prediction. Max diameter of the largest VOI ranked first at 9% relative importance, tied with NGTDM busyness at 9%. Total lesion glycolysis (TLG) contributed 8%, total metabolic tumor volume (TMTV) contributed 8%, and NGTDM coarseness contributed 5%. Notably, clinical features had much lower rankings - the highest-ranked clinical variable was R-IPI at position 12, with a relative importance of only 2.53%, suggesting that imaging features contain prognostic information that largely supersedes what clinical scoring systems capture.

TL;DR: The Dedicaid AutoML platform trained on Center 1 (n=41) using 100-fold Monte Carlo cross-validation with mixed ensemble learning. Top features by importance: max diameter (9%), NGTDM busyness (9%), TLG (8%), TMTV (8%), NGTDM coarseness (5%). R-IPI was the top clinical feature but ranked only 12th with 2.53% importance, indicating imaging features dominate over clinical scores.
Pages 6-7
Validation Results: 0.85 AUC on Independent External Dataset

The model's performance was evaluated at two levels. First, within-center performance was assessed by the 100-fold Monte Carlo cross-validation on the Center 1 dataset alone, yielding 66% sensitivity, 77% specificity, 78% positive predictive value (PPV), 70% negative predictive value (NPV), 71% accuracy, and an AUC of 0.74. These metrics represent a reasonable but modest baseline for a model trained and evaluated within the same institutional dataset.

Independent external validation on Center 2: When the model trained entirely on Center 1 data was applied to the independent Center 2 dataset (n=44), performance improved substantially: 79% sensitivity, 83% specificity, 69% PPV, 89% NPV, 82% accuracy, and an AUC of 0.85. The confusion matrix analysis tracked true positives, true negatives, false positives, and false negatives across all 44 Center 2 cases. The higher AUC on external validation (0.85) versus internal cross-validation (0.74) is an unusual finding - models typically perform worse on unseen external data than on cross-validated internal data.

Why external validation outperformed internal cross-validation: The authors provide a mechanistic explanation for this seemingly counterintuitive result. During the 100-fold Monte Carlo cross-validation within Center 1, each fold uses a subset of the 41 patients for training, further reducing the already small effective training set size and decreasing predictability. Additionally, random subsampling can create training-validation splits that are less representative of the true data distribution than two real-world clinical centers. Center 1 and Center 2 represent genuine clinical populations with natural similarity (same country, same standard treatment), whereas random splits within Center 1 create artificial variation. This interpretation suggests that the model's generalizability is strong and that its apparent internal performance underestimated its true predictive capacity.

Clinical interpretation of performance metrics: The 89% NPV achieved on the Center 2 external dataset is particularly clinically meaningful. An NPV of 89% means that when the model predicts a patient will achieve 2-year EFS (i.e., will not progress or relapse within 24 months), this prediction is correct 89% of the time. High NPV is valuable for identifying truly low-risk patients who may not require treatment intensification. The 79% sensitivity means the model correctly identifies 79% of patients who will actually have an event within 2 years, which is relevant for flagging high-risk patients who might benefit from closer monitoring or alternative treatment approaches.

TL;DR: Internal cross-validation (Center 1): AUC 0.74, sensitivity 66%, specificity 77%. Independent external validation (Center 2, n=44): AUC 0.85, sensitivity 79%, specificity 83%, NPV 89%, accuracy 82%. Unusually, external validation outperformed internal cross-validation - the authors attribute this to the small training subsets created by Monte Carlo splitting within Center 1.
Pages 7-8
What the Top Radiomic Features Actually Measure in DLBCL

Understanding what the top-ranked radiomic features represent biologically and metabolically provides insight into why they are prognostically informative in DLBCL. The two volume-based features - TMTV and TLG - have well-established prognostic roles in lymphoma literature. TMTV captures the total metabolically active tumor burden across all delineated lesions, and multiple prior studies have demonstrated that high TMTV at baseline predicts inferior progression-free and overall survival in DLBCL independent of R-IPI. TLG integrates both tumor volume and metabolic intensity, effectively capturing total glycolytic activity across the tumor burden.

Max diameter of the largest VOI: Interestingly, the maximum diameter of the single largest tumor focus ranked as the most important feature (9%), on par with NGTDM busyness. The authors note that within this cohort, max diameter of the largest lesion proved to be a better prognosis predictor than TMTV - a finding they attribute to population-level differences rather than a universal principle. They suggest that future investigations should examine whether max diameter or TMTV is more clinically relevant in different DLBCL populations, as these features, while correlated, may capture slightly different aspects of disease burden.

NGTDM busyness: The NGTDM (neighborhood gray-tone difference matrix) measures intensity transitions between neighboring voxels. Busyness specifically quantifies the spatial rate of intensity change - high busyness indicates rapid, non-smooth changes in FDG uptake between adjacent voxels, reflecting heterogeneous metabolic activity within the tumor microenvironment. In this study, Group 1 (poor prognosis) patients had higher busyness values. Biologically, the authors hypothesize that lymphoma cells embedded in necrotic or hypoxic periphery - regions where chemotherapeutic agents struggle to penetrate - contribute to this pattern of abrupt local intensity variation. High busyness may thus serve as an indirect marker of treatment-refractory intratumoral regions.

NGTDM coarseness: Coarseness, another NGTDM feature, captures large-scale spatial patterns in intensity. High coarseness indicates large homogeneous intensity regions within the lesion, while low coarseness indicates smaller, more fragmented texture subregions. Group 1 (poor prognosis) patients had lower coarseness values in this study, suggesting that tumors with more fragmented texture - potentially reflecting a more inhomogeneous mass with areas of necrosis, hypoxia, and variable cell proliferation rates - carry worse outcomes. The authors draw analogies to findings in other tumor types: coarseness has been highlighted as a prognostic feature in locally advanced rectal cancer, and both coarseness and busyness proved more predictive than SUVmax-based parameters in non-small cell lung cancer.

TL;DR: TMTV and TLG capture volumetric tumor burden, a well-validated prognostic signal. Max diameter was the single strongest predictor in this cohort. NGTDM busyness (high in poor-prognosis patients) reflects abrupt voxel-level intensity variation, potentially marking necrotic or hypoxic regions. NGTDM coarseness (low in poor-prognosis patients) suggests fragmented tumor texture associated with necrosis and heterogeneous proliferation.
Pages 8-9
The Relative Role of Clinical Variables Versus Imaging Features

One of the most clinically significant findings of this study is the comparison between imaging-derived features and standard clinical parameters for 2-year EFS prediction. Classical DLBCL prognostic tools like R-IPI rely on a small number of readily available variables: age, LDH level, performance status, Ann Arbor stage, and number of extranodal sites. While R-IPI has proven clinically useful for decades, it provides only coarse risk stratification and was not designed to capture the spatial metabolic complexity of the tumor.

R-IPI performance in this cohort: The chi-square statistical analysis confirmed that R-IPI, clinical stage, and cell-of-origin subtype (COO based on the Hans algorithm) were all significantly associated with 2-year EFS outcome (p = 0.015, 0.017, and 0.018, respectively). Patients with non-GCB subtype and higher R-IPI scores had significantly worse 2-year prognosis, consistent with extensive prior literature. However, when these same variables were included in the automated machine learning feature ranking, R-IPI emerged at position 12 with a relative importance of only 2.53% - well below the top imaging features.

Imaging as a surrogate for and superior to clinical parameters: The authors interpret this pattern as evidence that imaging features act as surrogates for - and in this context outperform - clinical parameters in predicting 2-year survival. This aligns with the broader literature showing that PET/CT-derived metrics can capture aspects of disease biology (metabolic burden, spatial heterogeneity) that are not reflected in clinical scores. The implication is that a purely imaging-based model, without requiring detailed laboratory workup or biopsy-based COO testing, could match or exceed the prognostic resolution of R-IPI.

Limitations of COO classification by IHC: The Hans algorithm used for COO classification in this study is an immunohistochemical proxy, not gene expression profiling. The Hans algorithm has moderate accuracy in replicating molecular COO classification - typically 75-90% concordance - meaning that some patients may have been misclassified into GCB versus non-GCB categories. If true molecular COO data had been available, the statistical associations with outcome and the feature importance rankings might have differed modestly. This is an inherent limitation of retrospective real-world cohort studies using IHC-based subtyping.

TL;DR: R-IPI, stage, and COO all significantly associated with 2-year EFS by chi-square (p approximately 0.015-0.018). However, in automated ML feature ranking, R-IPI placed only 12th at 2.53% importance - far below imaging features. The authors conclude imaging features are superior surrogates for prognosis. COO was based on the Hans algorithm (IHC proxy), not molecular gold-standard gene expression profiling.
Pages 9-11
Study Limitations and the Path Toward Clinical Deployment

Analysis of only the largest VOI: A key methodological limitation is that radiomic feature extraction was performed exclusively on the single largest tumor lesion per patient, rather than on all delineated lymphoma foci. DLBCL commonly presents with multifocal disease involving multiple nodal and extranodal sites, and the biology of smaller satellite lesions may differ from the dominant mass. The authors acknowledge this choice and justify it by noting that prior DLBCL radiomics studies have routinely analyzed only the largest VOI and still produced promising results, and that radiomic analysis of very small lesions is generally discouraged due to insufficient voxel counts for stable feature computation. Future work should explore whole-lesion radiomics across all PET-avid foci.

Small patient counts: With 41 patients in the training center and 44 in the validation center, this remains a relatively small dataset for machine learning. The class imbalance - approximately 2:1 ratio of remission to event patients - required automated handling within the AutoML pipeline. Small training sets are susceptible to overfitting and may not capture the full phenotypic diversity of DLBCL at the population level. The authors note, however, that using two different clinical centers with different PET/CT scanners strengthens the external validity of the findings relative to a larger but single-center study.

Retrospective design and selection bias: As a retrospective study, patient selection was constrained by the availability of pretreatment PET/CT scans performed within the defined time window (2014-2019), complete medical records, and receipt of standard R-CHOP-21 therapy. Patients treated with non-standard regimens were excluded, limiting generalizability to DLBCL patients managed outside the standard treatment paradigm. Prospective validation with consecutive patient enrollment would more reliably estimate real-world model performance.

Clinical implications and future directions: The authors conclude that predicting 2-year progression-free survival in DLBCL using imaging and clinical parameters is feasible with high-performance automated machine learning in a multi-center setting. The model's balanced sensitivity (79%) and specificity (83%) suggest it could serve as a decision-support tool to personalize treatment selection - for example, flagging high-risk patients for intensified initial therapy or early clinical trial enrollment. The study calls for larger prospective multi-center validation studies and emphasizes that the identified radiomic biomarkers (max diameter, NGTDM busyness, TLG, TMTV, NGTDM coarseness) warrant further investigation in diverse DLBCL populations and clinical settings before any clinical deployment.

TL;DR: Key limitations: radiomics restricted to the largest VOI per patient; small training (n=41) and validation (n=44) cohorts; retrospective design with patients treated only with standard R-CHOP-21. The dual-center design with different scanners partially offsets the small sample size by providing genuine external validation. The authors call for prospective multi-center studies to validate the five identified imaging biomarkers before clinical translation.