The International Prognostic Index in Aggressive B-Cell Lymphoma

Haematologica 2023 AI 7 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
A 30-Year-Old Model Still Defining Modern Lymphoma Care

The International Prognostic Index (IPI) was first published in 1993 by the International Non-Hodgkin's Lymphoma Prognostic Factors Project in the New England Journal of Medicine. It was developed from a dataset of patients with aggressive non-Hodgkin lymphoma who received anthracycline-containing combination chemotherapy on clinical trials between 1982 and 1987, a period when CD20 had only recently been identified as a protein, rituximab was still a decade from clinical availability, and pathology classifications relied on the International Working Formulation, Kiel, or Rappaport systems. The fact that this model, built from pre-rituximab era data, remains the standard prognostic and clinical trial stratification tool for aggressive B-cell lymphoma in 2023 is, as the authors describe, "frankly remarkable."

Context of the commentary: This 2023 landmark paper review, authored by Matthew J. Maurer of the Mayo Clinic, was published in Haematologica as part of a series revisiting pivotal studies in hematology. Maurer examines why the IPI has endured for three decades, what modifications have been attempted, how it performs in the current rituximab era, and what its continued use says about the challenges of developing durable prognostic tools in oncology.

The prognostic tool development challenge: Building a reliable, broadly applicable prognostic model requires assembling a dataset large enough for statistical modeling, representative of the real patient population, with sufficient follow-up to assess clinically meaningful endpoints such as overall survival and progression-free survival. Critically, a model can become clinically obsolete as treatment paradigms shift. The IPI has defied this obsolescence, retaining relevance across multiple treatment generations, though the reasons for this durability are themselves instructive.

The continued use of the IPI also reflects the challenge of improving on simple, well-validated clinical tools. While more computationally sophisticated models and molecular biomarkers have been developed, the combination of statistical simplicity, clinical accessibility, and three decades of validation data has made the IPI extraordinarily difficult to displace in routine practice.

TL;DR: The IPI was built from patients treated in 1982-1987, before rituximab existed, yet remains the standard prognostic tool for aggressive B-cell lymphoma in 2023. This commentary by Maurer (Mayo Clinic) examines why a pre-modern model has endured 30 years and what that longevity reveals about prognostic tool development in lymphoma.
Pages 1-2
Five Variables, One Score: How the IPI Works

The IPI is calculated from five binary clinical variables, each contributing one point to a total score ranging from 0 to 5. The five variables are: age greater than 60 years, Ann Arbor stage III or IV disease, Eastern Cooperative Oncology Group (ECOG) Performance Status of 2 or higher, lactate dehydrogenase (LDH) above the upper limit of normal, and involvement of two or more extranodal sites. Each variable is dichotomized, meaning patients either score a point or they do not, with no gradation for how far above normal a value falls or how many extranodal sites are involved beyond the two-site threshold.

Risk group categorization: IPI scores map to four conventional risk groups. Low risk (score 0-1) carries a five-year overall survival of approximately 73% in the original pre-rituximab dataset. Low-intermediate risk (score 2) corresponds to roughly 51% five-year survival. High-intermediate risk (score 3) gives approximately 43% survival, and high risk (score 4-5) the worst prognosis at around 26% five-year survival. In the rituximab era, these absolute survival figures have improved substantially across all groups, but the relative ordering and prognostic separation between groups has been maintained.

Why simplicity is a feature, not a limitation: From a statistical standpoint, dichotomizing continuous variables (such as LDH) discards information and reduces model efficiency. However, the authors note that this product-of-its-time simplicity has made the IPI extraordinarily clinician-friendly. The score can be calculated in real time at the bedside without electronic tools, reference charts, or laboratory interfaces beyond what is already ordered as part of standard staging workup. All five variables are routinely collected during initial evaluation of any patient with suspected aggressive lymphoma.

Scoring granularity: Notably, a score of 0 and a score of 5 represent extreme clinical profiles. A young patient with stage I disease, normal LDH, good performance status, and no extranodal involvement scores 0 and has an excellent prognosis. An older patient with stage IV disease, elevated LDH, poor performance status, and multiple extranodal sites scores 5 and has the worst prognosis. The model captures these extremes accurately, though its resolution in the middle range (scores 2-3) has been a target for refinement attempts.

TL;DR: The IPI sums five binary variables (age, stage, ECOG PS, LDH, extranodal sites) into a 0-5 score. Original dataset five-year overall survival ranged from 73% (score 0-1) to 26% (score 4-5). Dichotomization sacrifices statistical efficiency but produces a bedside-calculable tool requiring no electronic aids, which explains its durable clinical adoption.
Page 2
How the IPI Performs in the Modern R-CHOP Era

The IPI was derived from patients treated before the advent of rituximab, the anti-CD20 monoclonal antibody that transformed DLBCL outcomes when added to CHOP chemotherapy in the early 2000s. The addition of rituximab to CHOP (forming R-CHOP) improved three-year event-free survival from approximately 46% to 62% in the pivotal MInT trial, and improved overall survival across all IPI risk groups. This raised the fundamental question of whether a prognostic model trained on pre-rituximab outcomes would remain valid after this paradigm shift in therapy.

Preserved discrimination: Subsequent analyses, including the comprehensive comparison by Ruppert and colleagues published in Blood in 2020 examining IPI alongside the Revised IPI (R-IPI) and NCCN-IPI, confirmed that the IPI retains strong prognostic separation in the rituximab era. The original four-group IPI, the three-group R-IPI (collapsing low and low-intermediate into a "good prognosis" group and separating a "very good" group with score 0), and the NCCN-IPI (which incorporates age as a three-tier variable and distinguishes specific extranodal sites) all continue to stratify patients with meaningful discrimination.

R-IPI modification: The Revised IPI, proposed by Sehn and colleagues in 2007, recognized that with R-CHOP, the original IPI score of 0-1 low-risk group had become so favorable that further separation was warranted. The R-IPI identifies a "very good" group (score 0, four-year overall survival 94%), "good" (score 1-2, OS 79%), and "poor" (score 3-5, OS 55%). This regrouping improved practical utility without altering the underlying variables or their scoring.

NCCN-IPI refinements: The NCCN-IPI, developed from 1,650 patients treated at NCCN member institutions in the rituximab era, modified age weighting (giving additional points for age 41-60 and age over 75 compared to the original single cutoff at 60), adjusted LDH weighting, and identified specific high-risk extranodal sites (bone marrow, CNS, liver/GI tract, lung). These modifications improved discrimination, particularly at the extremes of the prognostic spectrum, but the magnitude of improvement over the original IPI has been modest in independent validation studies.

TL;DR: R-CHOP improved survival across all IPI groups but did not negate its prognostic separation. The R-IPI (Sehn 2007) identified a "very good" score-0 group with 94% four-year OS. NCCN-IPI (1,650 patients) refined age and LDH weighting and flagged high-risk extranodal sites, yielding modest improvements over the original. All three indices retain meaningful discrimination in the rituximab era.
Page 2
Efforts to Refine or Replace the IPI: What Has and Has Not Worked

The three decades since the IPI's publication have seen numerous attempts to improve on its prognostic performance by incorporating additional clinical, biological, imaging, and molecular variables. These efforts reflect an ongoing tension between model complexity and clinical utility: more variables can improve statistical discrimination, but each additional requirement reduces the tool's accessibility and applicability across different practice settings.

Biological and molecular additions: Numerous studies have investigated whether adding molecular markers improves on the IPI. Cell-of-origin classification (germinal center B-cell vs. activated B-cell subtype, determined by gene expression profiling or IHC surrogates such as the Hans algorithm) provides independent prognostic information but has not been incorporated into a universally adopted updated index. Double-hit and triple-hit lymphomas, defined by MYC rearrangements concurrent with BCL2 and/or BCL6 rearrangements, carry distinctly worse outcomes regardless of IPI score, but their identification requires fluorescence in situ hybridization (FISH) testing not available in all settings.

Imaging-based biomarkers: Total metabolic tumor volume (TMTV) measured from baseline PET/CT has emerged as a powerful independent prognostic biomarker in DLBCL. Multiple prospective cohort studies have shown that high TMTV (typically defined as greater than 220-300 cm3 depending on the measurement method) is associated with significantly inferior progression-free and overall survival, with hazard ratios in the range of 2-3. When combined with IPI, TMTV improves prognostic stratification, particularly by identifying a high-risk subgroup within the IPI low-risk category. However, TMTV measurement requires dedicated software, standardized PET protocols, and is not yet part of routine clinical reporting at most centers.

The marginal improvement problem: The IPI's persistent dominance partly reflects the difficulty of achieving substantial improvements in prognostic discrimination for a disease where treatment is standardized. When all patients receive similar frontline therapy, prognostic models can only separate those destined to respond from those who will not, and the biological determinants of treatment resistance are complex, heterogeneous, and not fully captured by any single set of clinical variables. The authors note that efforts to modify the IPI have made "marginal improvements" in prognostication, which explains why none of the proposed alternatives have fully displaced it in clinical practice or clinical trial stratification.

TL;DR: Molecular additions (COO subtype, double-hit status) and imaging biomarkers (TMTV, HR 2-3 for high vs. low) improve on IPI but face implementation barriers. Cell-of-origin adds prognostic information but requires gene expression profiling or validated IHC. Double-hit requires FISH. TMTV requires PET software. None has fully displaced the IPI in routine use or trial stratification because improvements have been statistically significant but clinically marginal relative to the IPI's simplicity.
Pages 1-2
The IPI as a Clinical Trial Stratification Standard

Beyond prognostication at the point of care, the IPI serves a critical function in the design and analysis of clinical trials for aggressive B-cell lymphoma. Nearly every major randomized trial in DLBCL over the past three decades, from the original MInT and GELA LNH98-5 trials establishing R-CHOP as standard of care, through the POLARIX trial establishing polatuzumab vedotin-based therapy, has used IPI score as a stratification variable in randomization and as an eligibility criterion for defining the target patient population.

Why consistent stratification matters: Clinical trials use prognostic stratification variables in randomization to ensure that treatment arms are balanced for known predictors of outcome. If one arm by chance enrolled more high-IPI patients, any observed treatment effect could reflect prognostic imbalance rather than true treatment benefit. Using the IPI as a stratification variable at randomization prevents this imbalance and also allows subgroup analyses by risk group to determine whether treatment effects are consistent across the prognostic spectrum.

Defining eligible populations: The IPI is also commonly used to define enrollment eligibility, either to enrich for high-risk patients (e.g., IPI score 2-5 or 3-5 in trials of intensified therapy) or to ensure a broadly representative population. Many investigational trials of novel frontline regimens require an IPI score of at least 2, which excludes the lowest-risk patients who already do well with standard R-CHOP. This deliberate enrichment for high-risk patients maximizes the statistical power of a trial to detect treatment benefit in patients who need it most.

Historical consistency: Because the IPI has been used as a stratification and eligibility criterion for 30 years, datasets from successive clinical trials can be compared against each other using IPI as a common reference point. This historical comparability would be lost if a new index were adopted, as results from trials stratified by a different tool would not be directly comparable to the extensive body of IPI-stratified data. This path dependency is a meaningful practical argument for the IPI's continued use even if a superior prognostic tool were developed.

TL;DR: The IPI has been the randomization stratification variable in nearly every major DLBCL trial for 30 years, including POLARIX. Its consistent use creates irreplaceable historical comparability across datasets. Trials commonly restrict enrollment to IPI 2-5 or 3-5 to enrich for high-risk patients. Replacing the IPI would break this longitudinal dataset consistency even if a new model were statistically superior.
Page 2
What Modern Computational Tools Would Change About Prognostic Modeling

The IPI was developed using the computational tools available in the late 1980s, which were substantially more limited than what is available today. The authors explicitly note this contrast: the smartphone apps, point-of-care calculation tools, modern data visualization platforms, and advanced modeling and data science techniques available in 2023 would have fundamentally different prognostic models if the IPI were being developed from scratch today rather than in 1993.

Modern modeling approaches: Contemporary prognostic model development for lymphoma would likely employ machine learning methods such as gradient-boosted decision trees (XGBoost, LightGBM), random forests, penalized regression (LASSO, elastic net), or deep neural networks trained on large integrated datasets. These approaches handle continuous variables without arbitrary dichotomization, can incorporate nonlinear relationships and interaction effects, and can integrate high-dimensional molecular, imaging, and clinical data simultaneously. Concordance index improvements of 5-15% over clinical models alone have been demonstrated in retrospective machine learning studies in DLBCL.

Electronic accessibility has changed the calculus: One of the IPI's key advantages in 1993 was bedside computability without electronic aids. In 2023, this advantage has largely disappeared: clinicians routinely use electronic health record interfaces, mobile apps, and online calculators for clinical decision support. The practical barrier to using a more complex model that requires a calculator has been substantially reduced. This shifts the cost-benefit analysis toward more sophisticated models that sacrifice mental arithmetic simplicity for improved prognostic accuracy.

From population to individual prognosis: A key limitation of all IPI variants is that they provide population-level risk estimates rather than individualized predictions. A patient with an IPI score of 3 belongs to a group with a certain average outcome, but significant heterogeneity exists within each IPI group. Molecular subtyping, genomic profiling, and multimodal AI models are beginning to decompose this within-group heterogeneity, moving toward more personalized prognostication. This transition parallels developments in other cancers where genomic signatures have refined prognostic stratification beyond what clinical indices alone can achieve.

TL;DR: The IPI's computational simplicity was a practical necessity in 1993 but is less relevant in 2023 when EHR calculators and mobile apps are ubiquitous. Modern ML methods (XGBoost, LASSO, neural networks) can handle continuous variables without dichotomization and incorporate molecular/imaging data, with C-index improvements of 5-15% over clinical models alone in retrospective DLBCL studies. However, population-level risk estimates from all IPI variants still cannot capture the within-group heterogeneity that molecular subtyping begins to address.
Page 2
The IPI's Boundaries and the Road Toward Better Prognostic Tools

The IPI's endurance reflects both its genuine strengths and the difficulty of replacing a well-validated tool with deep historical roots in clinical practice and clinical trial infrastructure. However, its limitations are real and become more apparent as the treatment landscape for aggressive B-cell lymphoma grows increasingly complex, with approved options now including polatuzumab vedotin-R-CHP, CAR-T cell therapy in the second-line setting (axicabtagene ciloleucel, lisocabtagene maraleucel), tafasitamab, loncastuximab tesirine, and bispecific antibodies such as epcoritamab and glofitamab.

Subtype-specific prognostication: The 2022 WHO classification and the concurrent International Consensus Classification have substantially expanded the recognized subtypes of large B-cell lymphoma, distinguishing entities such as high-grade B-cell lymphoma with MYC and BCL2 rearrangements (double-hit lymphoma), large B-cell lymphoma with IRF4 rearrangement, and primary mediastinal B-cell lymphoma. These entities have distinct biology, prognosis, and increasingly distinct optimal treatment approaches. A single IPI calculated identically across all of these entities obscures clinically meaningful differences that more subtype-specific prognostic models would capture.

Dynamic prognostication: The IPI is a static baseline assessment and cannot incorporate information that emerges during treatment, such as interim PET response (Deauville score), ctDNA clearance kinetics, or early-cycle pharmacokinetic data. Dynamic risk models that update prognosis as treatment proceeds are being developed and validated, but they require serial data collection and more complex computational infrastructure than the IPI demands. The integration of interim PET response into treatment decisions is already guideline-supported in Hodgkin lymphoma, and analogous adaptive approaches in DLBCL remain an active area of investigation.

The replacement challenge: Any candidate successor to the IPI faces a high bar: it must demonstrate superior prognostic discrimination in prospective multicenter validation, be implementable across academic and community settings globally, achieve regulatory acceptance for clinical trial use, and overcome the inertia of 30 years of IPI-based infrastructure. The authors suggest that the most likely path forward is not a single replacement index but a layered approach: the IPI as the universal baseline clinical tool, supplemented by molecular and imaging biomarkers in settings where they are available, and further refined by machine learning models as these mature toward prospective validation and clinical integration.

TL;DR: Expanding treatment options (CAR-T, bispecifics, ADCs) and refined WHO classification (double-hit lymphoma, IRF4-rearranged LBCL) increase the limitations of a single IPI applied uniformly. Dynamic models incorporating interim PET (Deauville) and ctDNA clearance are in development. The most plausible path forward is IPI as a universal baseline supplemented by molecular, imaging, and ML-based tools where available, rather than a single replacement index.