Diffuse large B-cell lymphoma (DLBCL) is the most common aggressive B-cell lymphoma, and its clinical behavior is highly variable. Approximately 60% of fit patients can be cured with upfront rituximab, cyclophosphamide, doxorubicin, vincristine, and prednisone (R-CHOP), but the remaining 30-40% develop relapsed or refractory disease associated with poor outcomes. This gap between early responders and treatment failures has driven a sustained effort to build more accurate prognostic models that can identify high-risk patients at the time of diagnosis, before they fail standard therapy.
Limitations of current clinical risk scores: Risk stratification in DLBCL has classically relied on the International Prognostic Index (IPI), the Revised IPI (R-IPI), and the NCCN-IPI, all of which aggregate five clinical variables: age, Ann Arbor stage, lactate dehydrogenase (LDH), Eastern Cooperative Oncology Group (ECOG) performance status, and number of extranodal sites. Recent benchmarking studies have shown concordance indexes (c-indexes) for overall survival in the range of 0.59-0.63 for these tools, reflecting their limited discriminative power. Notably, even the NCCN-IPI's highest-risk group carries a 5-year overall survival of approximately 49%, meaning a near-majority of "high-risk" patients survive long-term under standard treatment.
Molecular heterogeneity beyond the IPI: Gene expression profiling has identified three cell-of-origin (COO) subtypes with different survival trajectories under R-CHOP: germinal center B-cell-like (GCB), activated B-cell-like (ABC), and unclassified (UNC). A fourth high-risk group, molecular high-grade (MHG) lymphoma, is defined by a gene expression profile resembling double-hit or Burkitt lymphoma. Additionally, recurrent somatic mutation patterns have been used to define genomic subgroups - LymphGen assigns DLBCL tumors to biologically distinct clusters (BN2, EZB, MCD, N1, ST2, and others) each with characteristic mutation co-occurrence and clinical outcomes. Despite these advances, no single molecular classifier achieves sufficiently precise individual survival prediction for routine clinical use.
The LymForest-25 model: This study validates LymForest-25, a machine learning-based prognostic tool developed by the same group in prior work. LymForest-25 was built on 25 variables, comprising 19 gene expression transcripts, five IPI-component variables (age, LDH, Ann Arbor stage, extranodal sites, ECOG score), and the IPI score itself. The model uses a random forest framework and was originally trained on a Spanish cohort. The goal of this study is to test whether LymForest-25 retains its predictive accuracy when applied to an independent UK population-based cohort, and to explore whether integration with mutation-based classifiers such as LymphGen can push performance further.
Patient data for this external validation were drawn from the UK Haematological Malignancy Research Network (HMRN), a population-based registry. Gene expression data came from pre-treatment DLBCL biopsies of 644 R-CHOP-treated patients deposited in the Gene Expression Omnibus (accession: GSE181063), originally generated by Lacy et al. From this pool, 481 patients had complete annotation for all five IPI-component clinical variables and were included in the primary analysis. Gene expression was measured using Illumina HumanHT-12 WG-DASL V4.0 R2 bead chips, and expression values were rank-normalized before downstream modeling.
Gene coverage and preprocessing: LymForest-25 uses 19 gene expression transcripts as its molecular backbone. Of these, 17 were retrievable from the HMRN chip design: PSIP1, ADAM12, SGK196, BCL2A1, LMO2, SULF1, KLHL6, SNHG3, RAB3GAP2, HLA-DQB2, CPT1A, SCRG1, ATP8A1, LSMEM2/IFRD2, SLC5A12, FNBP1, and PDK1. The two missing genes - TRAV6 and FAM208B - were absent from the chip design and could not be imputed. This 17-gene partial signature was used in all models evaluated, and the authors explicitly acknowledge the omission as a study limitation.
Dimensionality reduction via PCA: Rather than using all 17 gene expression values as raw features in the Cox regression models, the authors applied principal component analysis (PCA) to the 17-gene expression matrix. They evaluated models using the full 17-gene matrix, the first principal component (PC1) alone, and incrementally larger sets (PC1+PC2, PC1+PC2+PC3, PC1+PC2+PC3+PC4, and five PCs). The four-component model (PC4) was identified as optimal based on AIC and bootstrapped c-index, striking the best balance between model complexity and predictive accuracy.
Mutation-based classification subgroup: A subset of 264 patients had mutation data available for LymphGen classification. Four mutation-based schemes were compared: a 6-cluster classification by Lacy et al., a modified 7-cluster version separating NOTCH1-mutated and MYC-rearranged cases, the standard LymphGen 9-cluster algorithm, and a modified LymphGen that adds an EZB-MYC group. These were tested for additive prognostic value over the LymForest-25 signature alone.
Statistical evaluation methods: Model performance was assessed using time-dependent area under the curve (AUC), Brier scores, Akaike information criterion (AIC), and Harrell's c-index. Bootstrapping without replacement with 500 cycles applied to 75% of the cohort, with 25% held out for cross-validation. Bootstrapped c-indexes were computed using the 632+ method. Optimism-adjusted c-indexes were also calculated to correct for overfitting. The Cox proportional hazards model from the R "survival" package was the primary modeling framework.
The first and central validation result of this study concerns whether the LymForest-25 gene expression signature, applied to the 17 available transcripts from the HMRN cohort, retains meaningful prognostic information. The answer is unambiguously yes. Time-dependent AUC analysis confirmed that the PC4 model - using the first four principal components of the 17-gene expression matrix - substantially outperformed the COO plus MHG classification at all four evaluated time points: 6 months, 1 year, 2 years, and 5 years after diagnosis.
Quantitative comparisons: At the 2-year time point, the LymForest-25 PC4 gene expression model alone achieved an AUC of 69.79%, compared to 64.14% for the COO+MHG classification. At 1 year, the gene expression model yielded an AUC of 69.12% versus 60.69% for COO+MHG. Brier scores, which measure the accuracy of probabilistic predictions (lower is better), also favored the LymForest-25 PC4 model consistently across all time points. The bootstrapped c-index and AIC both confirmed this hierarchy. Importantly, combining the PC4 gene expression signature with COO+MHG in a joint model provided no meaningful additional benefit over the PC4 signature alone, suggesting that the LymForest-25 transcripts already capture the prognostic information embedded in COO and MHG subtype assignments.
Which PCA model is best? Among the dimensionality reduction options tested, the model incorporating four principal components (PC4) consistently achieved the lowest AIC (2275) and the highest bootstrapped c-index (61.7). Adding a fifth component (PC5) slightly worsened performance (AIC 2277, c-index 61.5), indicating that the marginal information in the fifth component is outweighed by the increase in model complexity. The 17-gene raw matrix model achieved a c-index of 60.0 and AIC of 2290, performing meaningfully worse than the PC4 decomposition, which suggests that PCA-based dimensionality reduction captures the prognostic signal more efficiently than using raw gene expression values directly.
These results validate the core claim of LymForest-25: that a curated set of gene expression features, processed through a data-efficient representation, captures prognostic information that goes beyond what standard molecular subtype assignments can deliver. The fact that this holds with only 17 of the original 19 transcripts - and in an entirely independent UK population - strengthens confidence in the biological signal rather than cohort-specific overfitting.
While the LymForest-25 gene expression signature alone outperforms COO+MHG classification, the authors also evaluated whether integrating the PC4 signature with the IPI score further improves survival prediction. This is clinically important because the IPI is already routinely collected and calculated at diagnosis, making it a free addition to any molecular model. The results show that combining IPI with the LymForest-25 PC4 gene expression model provides the best overall prognostic accuracy of any configuration tested in this study.
Performance numbers for the combined model: The IPI score combined with the LymForest-25 PC4 signature achieved AUC values of 74.56% (6 months), 77.48% (1 year), 78.73% (2 years), and 76.45% (5 years). For comparison, the IPI score alone achieved AUC values of 74.70% (6 months), 74.27% (1 year), 74.57% (2 years), and 72.83% (5 years). The addition of the gene expression signature therefore adds approximately 3-4 percentage points of AUC at the 1-year and 2-year time points, with minimal change at 6 months and modest improvement at 5 years. Brier scores also favored the combined model, with the IPI+PC4 model achieving 0.058 at 6 months and 0.124 at 2 years compared to 0.061 and 0.136, respectively, for the IPI alone.
COO+MHG adds nothing to the IPI either: The IPI combined with COO+MHG achieved AUC values of 70.38% (6 months), 74.11% (1 year), 76.98% (2 years), and 75.95% (5 years), which are comparable to or slightly below those of IPI+PC4 at most time points. Adding the LymForest-25 PC4 signature on top of both IPI and COO+MHG yielded AUC values nearly identical to IPI+PC4 alone (72.05%, 75.91%, 78.57%, 76.96%), again confirming that the gene expression signature largely subsumes the prognostic information in COO and MHG assignments.
The practical implication is that clinicians seeking the best available survival estimate at diagnosis should use LymForest-25 combined with the IPI, rather than COO+MHG combined with the IPI. The former produces meaningfully better 1-year and 2-year predictions with the same clinical data input, and it does not require a separate molecular subtype determination that the gene signature already supersedes.
Beyond overall cohort performance, the study evaluated whether the LymForest-25 PC4 gene expression signature retains prognostic value within each COO subgroup and within the MHG group separately. This analysis is important because a true molecular prognostic tool should be informative even after patients are partitioned by a first-level classifier - it should stratify risk within GCB, within ABC, and within the "unclassified" subgroup, rather than merely reflecting the known differences between these broad categories.
Unclassified COO subgroup shows strongest performance: The LymForest-25 PC4 signature performed best within the unclassified COO subgroup, achieving time-dependent AUC values above 70% at all evaluated time points: 74.41% at 6 months, 78.47% at 1 year, 78.81% at 2 years, and 73.69% at 5 years. This is a particularly notable finding because the unclassified subgroup - comprising approximately 17.7% of the cohort - is precisely where conventional classification offers the least prognostic guidance. Patients labeled "unclassified" by COO are by definition those whose gene expression does not clearly align with either GCB or ABC profiles, and yet LymForest-25 identifies high and low-risk strata within this group with substantial accuracy.
ABC and GCB subgroups: Within the ABC subgroup (27.9% of the cohort), AUC values ranged from 64.40% at 6 months down to 53.21% at 5 years, with better prediction of early events than late ones. Within the GCB subgroup (48.0% of the cohort), AUC ranged from 61.42% at 6 months to 55.81% at 5 years. Both subgroups show meaningful but more modest performance compared to the unclassified group, with a consistent pattern of better prediction at early versus late time points - suggesting that short-term mortality risk is more molecularly tractable within these groups than long-term survival differences.
MHG subgroup: Within the MHG group (6.4% of cohort, N=31), the signature showed a somewhat reversed pattern: less accurate prediction of early events (53.19% at 6 months, 49.48% at 1 year) but better performance for later time points (58.22% at 2 years, 56.04% at 5 years). The authors caution that the small MHG sample size (N=31) limits the reliability of these estimates and precludes strong conclusions. The overall message for MHG is that this very high-risk group has high baseline mortality that is difficult to further stratify within, rather than the signature being uninformative per se.
One of the most clinically significant findings in this paper is the age-dependent performance of molecular prognostic models. The cohort was split into younger patients (age 70 years or younger, N=302) and older patients (over 70 years, N=179), and all major model configurations were evaluated separately in each group. The results reveal a consistent and important asymmetry: molecular predictors add substantial value over the IPI in younger patients but provide essentially no improvement over the IPI alone in older patients.
Younger patients - strong molecular signal: In patients 70 years or younger, the IPI + LymForest-25 PC4 model achieved AUC values of 81.93% at 6 months, 78.06% at 1 year, 81.30% at 2 years, and 77.80% at 5 years. The IPI score alone in this subgroup achieved 78.00%, 73.93%, 76.11%, and 74.06% at the same time points. The gene expression signature therefore adds approximately 4-5 AUC percentage points consistently in younger patients. Brier scores follow the same pattern, with the combined model achieving 0.039 at 6 months and 0.100 at 2 years versus 0.042 and 0.110 for IPI alone.
Older patients - molecular predictors plateau: In patients older than 70 years, the IPI score alone achieved AUC values of 68.75% (6 months), 73.43% (1 year), 70.93% (2 years), and 68.50% (5 years). Adding the LymForest-25 PC4 gene expression signature to the IPI in this age group produced virtually identical AUC values: 60.30% (6 months), 70.01% (1 year), 69.56% (2 years), and 68.96% (5 years) - with no meaningful improvement and even slight degradation at the 6-month time point. Brier scores similarly showed no improvement from adding molecular data in this age group.
Why do molecular models lose power in older patients? The authors propose that this limitation reflects the increasing influence of non-disease factors on survival in older DLBCL patients. Comorbidities, frailty, and treatment dose intensity - variables not captured in gene expression profiles or the standard IPI components - become progressively more important determinants of outcome as age increases. The authors suggest that incorporating geriatric assessment scores into prognostic models for older DLBCL patients could recover some of the predictive performance that molecular variables cannot provide.
The final analytical arm of this study examines whether mutation-based DLBCL classification systems add prognostic value on top of the LymForest-25 gene expression model. This analysis was restricted to the 264-patient subgroup with mutation annotation available, and it tested four mutation classification schemes: the Lacy et al. 6-cluster model, the modified Lacy et al. 7-cluster model, the standard LymphGen 9-cluster algorithm, and a modified LymphGen classification that adds an EZB-MYC group to capture tumors with both EZB cluster membership and MYC rearrangement.
LymphGen cluster distribution in this cohort: Within the 264-patient mutation subgroup, LymphGen classification assigned patients as follows: BN2 (9.85%), EZB (23.86%), MCD (7.95%), N1 (3.03%), ST2 (7.58%), BN2/N1 (0.38%), EZB/ST2 (1.14%), MCD/ST2 (0.38%), and unclassified (45.83%). The high unclassified rate - nearly half of all mutation-characterized patients - reflects a known limitation of the LymphGen algorithm, which achieves confident classification in only a subset of DLBCL cases using standard mutation panel data.
Statistical approach and constraints: The very small cluster sizes in some mutation categories (e.g., N1 at 3.03%, BN2/N1 at 0.38%) precluded bootstrapping for time-dependent AUC and Brier score comparisons - the standard validation approach used for the full cohort analysis. Instead, the authors relied on AIC and optimism-adjusted c-indexes, which are less robust metrics that carry greater uncertainty. This is an important methodological caveat when interpreting the mutation integration results.
Results - modified LymphGen shows a modest signal: For c-indexes, no relevant improvement was observed when any of the four mutation-based classifications was added to LymForest-25. For AIC, LymForest-25 alone was superior to its combinations with the standard Lacy et al. models and the standard LymphGen classification (higher AIC = worse fit). However, the modified LymphGen model - which adds the EZB-MYC group - produced a difference of more than 2 AIC units when combined with LymForest-25, which is the conventional threshold for considering an improvement meaningful. This suggests a modest but statistically notable benefit from incorporating the EZB-MYC hybrid genetic subgroup alongside the gene expression model, though the authors appropriately note that these results require prospective confirmation given the analytical constraints of the subgroup.
Missing transcripts: The most concrete methodological limitation is the unavailability of 2 of the 19 LymForest-25 transcripts (TRAV6 and FAM208B) from the HMRN gene expression chips. While the 17-gene partial model still outperformed COO+MHG classification significantly, it is possible that incorporating all 19 transcripts would yield even better performance. The authors cannot quantify this potential improvement from the current analysis. This limitation also highlights a practical barrier to clinical translation: prognostic gene expression models require the specific platform and probe coverage used in their development, limiting portability across laboratories using different expression measurement technologies.
Bootstrapping rather than true external validation: Although this study presents itself as an external validation, the 481-patient cohort is treated as a single dataset that is internally partitioned using bootstrapping (75% training, 25% test, 500 cycles). This is stronger than simple split-validation but is not the same as testing a completely locked, pre-specified model on an entirely separate dataset collected prospectively. The bootstrapped c-indexes and AIC comparisons should therefore be interpreted as well-estimated internal validation metrics rather than fully independent external validation - a distinction the authors acknowledge.
Older patients and the frailty gap: As discussed in the age-stratified analysis, the failure of molecular predictors to improve IPI performance in patients older than 70 years points to an important unmet need. The authors suggest that incorporating standardized geriatric assessment scores - which capture frailty, comorbidities, and functional status more accurately than age and ECOG alone - could improve prognostic models for older DLBCL patients. Several groups have validated simplified frailty scores in older DLBCL patients with promising results.
Context within the broader field: The authors note that their work aligns with other machine learning approaches to lymphoma prognostication, citing a contemporaneous study by Zaccaria et al. that used unsupervised clustering on baseline clinical data from a mantle cell lymphoma phase III trial. That model also validated across independent cohorts and similarly showed reduced performance in older patients. The convergence of findings across studies suggests that the age boundary for molecular predictor utility is a real biological and clinical phenomenon rather than a study-specific artifact. Future directions include integrating geriatric assessment, expanding validation to additional independent cohorts, and developing platform-agnostic implementations of the LymForest-25 gene signature that do not require a specific microarray chip design.