Remaining Challenges in Predicting Patient Outcomes for Diffuse Large B-Cell Lymphoma

Expert Review of Hematology 2019 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Why Predicting DLBCL Outcomes Remains Difficult

Diffuse large B-cell lymphoma (DLBCL) is the most common type of non-Hodgkin lymphoma and is fatal without treatment. It is characterized by significant clinical and molecular heterogeneity, meaning that patients with the same diagnosis can have vastly different disease biology, treatment responses, and survival trajectories. Over the past three decades, researchers have developed a diverse array of prognostic tools drawing on clinical factors, cell-of-origin (COO) subtypes, and genomic subgroups, but no single method has emerged as the definitive standard for predicting outcomes.

The International Prognostic Index (IPI): Developed more than 25 years ago using stepwise regression, the IPI remains the most widely used clinical risk assessment tool for DLBCL. It incorporates five factors: age, disease stage, serum lactate dehydrogenase (LDH) level, performance status, and number of extranodal disease sites. However, survival outcomes shifted markedly after the addition of rituximab to frontline chemotherapy regimens, exposing limitations of the original IPI. Two variants have been proposed: the revised-IPI (R-IPI), which re-groups the original scores into three risk categories, and the NCCN-IPI, which assigns incremental scores to increasing levels of age and LDH and includes specific high-risk extranodal sites.

Head-to-head comparison: A pooled analysis of individual patient-level data from 7 multicenter trials (n=2,561, 86% DLBCL) treated with R-CHOP or variants found that the NCCN-IPI produced the greatest absolute difference in overall survival (OS) estimates between the highest and lowest risk groups at 1, 3, and 5 years. The NCCN-IPI best discriminated OS with a c-index of 0.63, but this was only marginally better than the original IPI (c-index = 0.62). These modest c-index values underscore how much room remains for improvement in clinical risk stratification.

This review by Harkins et al. from Emory University and Winship Cancer Institute surveys the full landscape of DLBCL prognostic methods, including clinical scoring systems, molecular classification, genomic sequencing, sociodemographic factors, treatment response measures, and machine learning approaches, identifying where each falls short and where integration of multiple data types may finally yield accurate, personalized predictions.

TL;DR: DLBCL is the most common non-Hodgkin lymphoma with highly heterogeneous outcomes. The standard IPI and its variants (R-IPI, NCCN-IPI) achieve only modest discrimination (c-index 0.62-0.63). This review surveys clinical, molecular, genomic, sociodemographic, and machine learning approaches to prognostication.
Pages 3-4
Beyond the IPI: Fitness Status, Hematologic Parameters, and Vitamin D

Machine learning-enhanced clinical models: Biccler et al. developed a new prognostic model using a stacking algorithm, an ensemble learning method that aggregates several regression models to generate survival curves without requiring specification of a single modeling approach. Trained on clinical data from Danish and Swedish nationwide lymphoma registries, this stacking-based model was reported to be superior to both the NCCN-IPI and IPI (available at lymphomapredictor.org). Separately, Howlader et al. used the large SEER registry to build a model incorporating non-IPI risk factors such as gender, race, Hispanic ethnicity, marital status, and a population-level measure of poverty, estimating 10-year survival and cure rates across risk groups.

Comprehensive Geriatric Assessment (CGA): The CGA categorizes patients as fit, unfit, or frail and incorporates performance status, functional status, comorbidities, nutritional status, cognitive function, psychological state, social support, and polypharmacy. Prospective studies have shown that CGA stratifies older DLBCL patients by OS, overall response rate (ORR), and treatment-related toxicities. CGA was more effective than clinical judgment alone in identifying fit patients who could tolerate anthracycline-based treatment with curative intent. However, the full CGA is impractical in routine clinical settings, and simplified models still require validation.

Lymphocyte-to-monocyte ratio: The ratio of lymphocytes to monocytes in serum at DLBCL diagnosis may reflect the immune microenvironment and can predict response to R-CHOP independently of IPI score. While these indicators of tumor milieu may help predict disease progression, molecular analysis of the microenvironment will be essential for developing therapies that target the underlying drivers of poor-risk microenvironments.

Vitamin D: Low 25-hydroxyvitamin D levels are associated with inferior DLBCL outcomes. A prospective study of 983 newly diagnosed NHL patients found that 25(OH)D insufficiency (< 25 ng/mL) was associated with inferior event-free survival (EFS) and OS. In a cohort of 155 aggressive lymphoma patients treated with R-CHOP, two-thirds were vitamin D deficient (< 20 ng/mL), and normalization of 25(OH)D following weekly 25,000 IU cholecalciferol supplementation led to improved EFS compared to patients with persistent deficiency.

TL;DR: Beyond the IPI, prognostic factors include machine learning-based stacking models (superior to NCCN-IPI), Comprehensive Geriatric Assessment for older patients, lymphocyte-to-monocyte ratio as a microenvironment marker, and vitamin D status (deficiency in 66% of patients, correctable with supplementation improving EFS).
Pages 4-6
Cell-of-Origin Subtyping and IHC vs. Gene Expression Profiling

Gene expression profiling (GEP) identified two molecularly distinct DLBCL subgroups: germinal center B cell-like (GCB) and activated B cell-like (ABC). Despite identical histologic appearance, ABC-DLBCL has significantly worse outcomes, with a 5-year OS rate of 35% after anthracycline-based chemotherapy compared to 60% for GCB-DLBCL. Understanding which COO subtype a patient has is therefore critical for both prognosis and the design of subtype-specific therapies.

Immunohistochemical (IHC) surrogates: Because GEP was historically limited by cost, turnaround time, and reliance on fresh frozen tissue, IHC classification systems were developed as practical surrogates. The Hans algorithm, the most widely used, classifies GCB and non-GCB phenotypes using CD10, BCL6, and MUM1 expression with 86% concordance with GEP, meaning 14% of patients are misclassified. Eight additional IHC algorithms have been developed (including Choi, which adds GCET1 and FOXP1, and Tally, which scores CD10, MUM1, GCET1, FOXP1, and LMNO2), but none have achieved widespread clinical adoption. Critically, multiple studies found poor inter-assay and inter-laboratory concordance for IHC methods, and many failed to reproduce the prognostic significance of COO subtyping seen with GEP.

GEP advances and Lymph2Cx: The shift from fresh frozen tissue to formalin-fixed paraffin-embedded tissue (FFPET) with 98% concordance made GEP more clinically feasible. The Lymph2Cx assay, developed by the Lymphoma/Leukemia Molecular Profiling Project using the NanoString platform, can determine COO subtype from FFPET biopsies with both consistency and accuracy. In 2019, Robetorye et al. published clinical validation of Lymph2Cx in routine practice, analyzing 90 clinical cases and reporting it was accurate, rapid, and reproducible. However, barriers remain: cost-effectiveness analysis has not been conducted, sufficient tissue beyond diagnostic workup is required, and access to specialized equipment is needed.

Digital pathology and deep learning: Informatics approaches offer the potential for more objective classification. Survival convolutional neural networks combine deep learning with traditional survival models to learn survival-related patterns from digitized H&E and IHC whole-slide images. These networks include convolutional layers for visual pattern extraction and a Cox proportional hazards layer for modeling time-to-event data. The authors' group at Emory has begun applying these image analysis tools with machine learning to classify COO for DLBCL consistently and accurately.

TL;DR: ABC-DLBCL has 5-year OS of 35% vs. 60% for GCB-DLBCL. The Hans IHC algorithm achieves 86% concordance with GEP (14% misclassification). Lymph2Cx on the NanoString platform offers accurate FFPET-based COO subtyping. Deep learning survival CNNs are emerging as tools for more objective, image-based classification.
Pages 7-9
Next-Generation Sequencing and Novel Genomic Subtypes

Next-generation sequencing has revealed the driving genetic aberrations in DLBCL and their preferential distribution among subtypes. ABC-DLBCLs preferentially harbor somatic mutations in CD79A/B and MYD88 that result in constitutive B-cell receptor (BCR) signaling and NF-kB activation, while GCB-DLBCLs have much higher rates of mutated EZH2, which suppresses apoptosis. These molecular hallmarks are increasingly accepted as nascent biomarkers for selecting precision medicine therapies.

Four landmark genomic studies: Reddy et al. performed whole-exome and transcriptome sequencing of 1,001 DLBCL patients and identified 150 driver mutations, building a genomic risk model that was highly predictive, especially for long-term mortality outcomes where the clinical IPI model fell short. Schmitz et al. sequenced 574 DLBCL samples and defined four genomic subtypes beyond COO: MCD (MYD88 and CD79B mutations), BN2 (BCL6 fusions and NOTCH2 mutations), N1 (NOTCH1 mutations), and EZB (EZH2 mutations and BCL2 translocations). Chapuy et al. analyzed 304 primary DLBCLs and identified 98 driver mutations defining 5 novel subsets. Arthur et al. studied 338 de novo cases and identified novel cis-regulatory sites and recurrent 3' UTR mutations in NFKBIZ as NF-kB pathway activators in ABC-DLBCL.

Double-hit and double-expressor lymphomas: "Double-hit" lymphomas (DHL) with dual rearrangement of MYC and either BCL2 or BCL6 are associated with poor outcomes after standard R-CHOP. Ennishi et al. developed the DHITsig, a 104-gene expression signature that identifies a GCB subgroup with distinct biology: 27% of GCB-DLBCLs were DHITsig-positive, and only half of those exhibited actual MYC and BCL2 rearrangements by FISH. The DHITsig-negative GCB subgroup exhibited a 5-year disease-specific survival of 90%, suggesting R-CHOP alone is sufficient for those patients. Separately, double-expressor lymphomas (DELs) with high IHC expression of both BCL2 and MYC carry up to a 9-fold elevated risk of death, but pathologist scoring is inconsistent: BCL2 IHC agreement was just 47% (kappa = 0.23) and MYC scoring yielded nearly 40% discordant cases.

Oxidative phosphorylation subgroup: The OxPhos-DLBCL subgroup exhibits a gene expression profile independent of COO classification. These lymphomas do not respond to BCR signaling inhibition but may respond to disruption of fatty acid oxidation, glutathione synthesis, and PPAR-gamma. The antibiotic tigecycline exhibited toxicity in OxPhos-DLBCLs at doses tolerable in humans, suggesting potential therapeutic avenues.

TL;DR: Four major sequencing studies (574-1,001 patients each) defined novel DLBCL genomic subtypes (MCD, BN2, N1, EZB) and identified 98-150 driver mutations. The DHITsig captures twice as many high-risk GCB patients as FISH alone. BCL2 IHC scoring agreement is only 47% (kappa = 0.23), highlighting the need for standardized molecular methods.
Pages 9-10
Translating High-Dimensional Genomic Data into Clinical Predictions

The integration of gene expression data with clinical, histological, imaging, demographic, and epidemiological information holds promise for improving cancer prognosis, but the enormity and complexity of the resulting datasets present formidable analytical challenges. Machine learning methods are specifically designed to organize, process, and discover actionable knowledge in these high-dimensional settings, addressing three fundamental predictive tasks: cancer susceptibility (risk assessment), cancer recurrence, and cancer survival prediction.

Dimensionality reduction: Two key techniques are used to handle the curse of dimensionality in genomic data. Feature selection methods (filtering based on relevance scores, stepwise selection, Lasso regression) identify the most informative variables, producing more compact and interpretable models while avoiding overfitting. Feature extraction methods (hierarchical clustering, principal component analysis) create new composite variables. An important drawback of unsupervised approaches like clustering and PCA is that the new features are not guaranteed to be predictive of clinical outcome. Supervised PCA, which selects principal components based on clinical outcome, addresses this limitation and has been used to predict patient survival from microarray data.

Key DLBCL studies: Shipp et al. combined gene expression data from 58 DLBCL patients with classification analysis to predict cured versus fatal/refractory disease, though the analysis was limited by low gene count and missing clinical data elements. More recently, Reddy et al. performed an integrative analysis of whole-exome and transcriptome sequencing in 1,001 DLBCL patients, using a supervised learning approach with embedded feature selection to develop a predictive model incorporating clinical information, 150 genetic driver genes, and gene expression markers (COO, MYC, BCL2). The authors note that future efforts should expand beyond linear models, comprehensively assess all combinations of features and classification methods, and use alternative feature selection algorithms capable of capturing DLBCL's molecular heterogeneity.

Deep learning and the CIRI model: Neural network-based approaches (deep learning) have shown promise for precision genomic medicine due to their ability to leverage large datasets for clinical outcome prediction. The Continuous Individualized Risk Index (CIRI) dynamically integrates personalized risk factors at pretreatment, treatment, and end-of-treatment phases, showing improved outcomes prediction compared to static models. The authors argue for unified machine learning platforms that can incorporate imaging, histological, clinical, and genomic data simultaneously.

TL;DR: Machine learning methods including Lasso regression, supervised PCA, ensemble stacking, and deep learning are being applied to DLBCL prognostication. Reddy et al. built a model from 1,001 patients integrating 150 driver genes with clinical data. The CIRI model dynamically integrates risk factors across treatment phases, outperforming static models.
Pages 11-12
Race, Socioeconomic Status, and Place of Residence in DLBCL Outcomes

Racial disparities: SEER data have shown that African American DLBCL patients present approximately 10 years younger, exhibit more advanced stage at diagnosis, and have inferior 5-year survival relative to white patients. These differences have historically been attributed to social, environmental, and behavioral factors, but biologic factors may also play a role. Lee et al. used genetically-determined African ancestry rather than self-reported race to examine differences in mutational profiles of 150 DLBCL driver genes. They found distinct prevalence and patterns of genomic alterations in African American patients, suggesting involvement of different oncogenic genes and pathways that may contribute to racial differences in disease incidence, presentation patterns, and survival.

Socioeconomic status (SES): A retrospective cohort analysis of 33,032 DLBCL patients diagnosed in California from 1988 to 2009 found that patients living in lower-SES neighborhoods had increased mortality risk compared to higher-SES neighborhoods, even after adjusting for insurance status. Using the National Cancer Database (NCDB), uninsured patients (hazard ratio 1.39, p<0.05) and Medicaid-insured patients (HR 1.48, p<0.05) had lower survival than privately insured patients after adjusting for age, sex, race, ZIP code area, and education level.

Place of residence: Analysis of 83,108 DLBCL patients in the NCDB found that rural and urban patients were more likely than metro populations to have lower SES, Medicaid insurance, advanced stage at diagnosis, and more comorbidities. Rural and urban populations exhibited inferior 5-year OS compared to metro patients, although risk was attenuated by SES, insurance status, and treatment facility type. Neighborhood SES may affect outcomes through mechanisms including healthcare availability, access to healthy foods, environmental pollution, health literacy, and social support.

Diagnosis-to-treatment interval (DTI): Analysis of 986 DLBCL patients from the Molecular Epidemiology Resource database (Mayo Clinic and University of Iowa, 2002-2013) found that shorter DTI was associated with adverse clinical factors and worse outcomes. The prognostic effect of DTI independent of IPI may indicate high-risk disease features not captured by standard prognostic assessments. The LEO Cohort Study, accruing at 8 U.S. medical centers with a goal enrollment of 13,900 patients, aims to catalog the interplay of clinical, epidemiologic, genetic, and treatment factors, including quality-of-life scores and NHL molecular subtype not available in SEER.

TL;DR: African American DLBCL patients present 10 years younger with more advanced disease. In a 33,032-patient California cohort, lower-SES neighborhood residence independently predicted higher mortality. Uninsured (HR 1.39) and Medicaid-insured (HR 1.48) patients had worse survival. The LEO Cohort (goal: 13,900 patients) is the first large U.S. NHL study to capture molecular subtype alongside sociodemographic factors.
Pages 13-14
Interim PET/CT and Circulating Tumor DNA for Dynamic Outcome Prediction

Interim PET/CT: While pretreatment PET/CT staging is standard, the use of interim PET/CT (between cycles 2 and 4 of frontline therapy) offers the possibility of adapting treatment based on ongoing response. This strategy has produced improved outcomes in Hodgkin lymphoma and is standard of care for that disease. In DLBCL, however, results have been variable. A retrospective study from the authors' group showed that interim PET with full resolution of metabolic tumor was highly correlated with achieving complete remission and improved survival. However, the degree of decrease from pre-treatment to interim PET did not relate to outcomes. Pretreatment total metabolic tumor volume on FDG PET/CT has independently been shown to predict OS, and baseline tumoral metabolic heterogeneity may be prognostic.

Circulating tumor DNA (ctDNA): Cell-free DNA fragments in plasma, mostly originating from the tumor, can be detected and sequenced thanks to improvements in DNA sequencing technology. ctDNA is attractive because it is minimally invasive, allows serial sampling, and can detect subclinical disease. Armand et al. first showed that ctDNA is detectable in newly diagnosed DLBCL and becomes undetectable after treatment. Roschewski et al. found that detection of VDJ segments of tumor immunoglobulin genes after 2 treatment cycles correlated with disease progression by 5 years, with ctDNA detection occurring a median of 3.5 months before clinical evidence of disease in surveillance patients. Kurtz et al. confirmed prospectively that molecular disease in plasma preceded PET/CT detection of relapsed disease.

Tumor heterogeneity and clonal evolution: Rossi et al. demonstrated that ctDNA can detect mutations undetectable in tissue biopsy (due to spatial tumor heterogeneity) and can identify new mutations in treatment-resistant patients, potentially reflecting resistance mechanisms. ctDNA can also be used for genotyping to identify COO and to distinguish patterns of clonal evolution in transformed versus indolent lymphomas.

Both interim PET/CT and ctDNA offer longitudinal measures of therapeutic efficacy, but additional studies are needed to determine whether routine ctDNA surveillance is cost-effective and improves clinical outcomes before widespread adoption.

TL;DR: Interim PET/CT with full metabolic resolution correlates with complete remission in DLBCL, though the approach has variable results across studies. ctDNA detects disease progression a median of 3.5 months before clinical evidence and can reveal tumor heterogeneity and resistance mutations invisible to tissue biopsy.
Pages 15-16
Toward Integrated, Multi-Factor Prognostic Models

Genomics-driven precision medicine: The authors envision that within 5 years, advances in sequencing technology combined with robust population-level capture of clinical and sociodemographic factors will allow widespread, real-time incorporation of complex genomic and patient-specific data into prognostic models. Large genomic studies using FFPE blocks at diagnosis and ctDNA at multiple timepoints (diagnosis, between cycles, end of treatment) will be instrumental in defining rational biomarkers for therapy selection. Standardized sequencing methods and informatics pipelines across multi-center collaborations will be essential for reproducibility.

Bioinformatics and pathway cross-talk: Novel computational approaches to define cross-talk between molecular pathways may identify relevant genes or gene groups that serve as biomarkers for therapeutic response to targeted agents. Quick turnaround of genomic analysis will be critical given DLBCL's aggressive nature and the frequent need for urgent treatment, as patients with the shortest diagnosis-to-treatment intervals have the worst outcomes and may have the greatest need for molecularly-targeted approaches.

Clinical workflow integration: Converting "big data" into a streamlined, clinically relevant report integrated into the electronic medical record will be essential for clinician adoption. This requires collaborative relationships between clinicians, bioinformaticians, and other professionals. A systematic method for balancing clinical and biological prognostic factors when determining treatment strategies by subtype or patient must be developed. Ensuring rapid and affordable access to genomic sequencing is necessary for confirming whether patients express biomarkers linked to specific treatment strategies in a cost-effective manner.

Addressing disparities: The advances in molecular technology described in this review risk exacerbating existing disparities if these technologies remain inaccessible to certain patient groups. The authors emphasize that eliminating disparities requires understanding the interaction between biological, clinical, and socioeconomic factors, alongside analysis of community infrastructures (transportation, sick pay) necessary to reduce barriers to care. Public policy informed by practice pattern analysis and health outcomes research will be needed to improve access for poor-risk populations.

TL;DR: The future of DLBCL prognostication lies in integrating genomic, clinical, and sociodemographic data through machine learning platforms embedded in clinical workflows. Key challenges include standardizing multi-center sequencing pipelines, achieving rapid turnaround for an aggressive disease, and ensuring equitable access to molecular technologies to avoid widening existing disparities.