Diffuse large B-cell lymphoma (DLBCL) is the most common subtype of lymphoma and accounts for a substantial share of NHL diagnoses worldwide. Its defining clinical feature is heterogeneity: more than 60% of patients can be cured with standard rituximab, cyclophosphamide, doxorubicin, vincristine, and prednisone (R-CHOP) therapy, but a meaningful fraction relapse or are refractory from the outset. The central challenge is predicting in advance which patients fall into which category, so that non-responders can be redirected toward alternative regimens or clinical trials rather than being exposed to ineffective treatment.
Existing molecular classifications: Multiple classification frameworks have been developed over the past two decades in an effort to stratify DLBCL biologically. The earliest approach used microarray-based gene expression profiling to divide DLBCL into two cell-of-origin (COO) groups: germinal center B-cell-like (GCB) and activated B-cell-like (ABC), with ABC carrying worse prognosis under R-CHOP. The GenClass algorithm later added genetic abnormalities (MCD, BN2, N1, EZB subgroups), but could classify only 54% of DLBCL cases. The extended LymphGen algorithm expanded this to seven subgroups and improved coverage, while Chapuy et al. used mutation profiling and chromosomal structural abnormalities to define five subgroups. FISH-based double-hit and triple-hit classifications (rearrangements in MYC co-occurring with BCL2, BCL6, or both) identify an especially aggressive subset for which R-CHOP is ineffective.
The clinical gap: Despite this proliferation of biological subclassification schemes, none reliably predicts overall survival or progression-free survival at the individual patient level. Critically, many require whole-exome sequencing or complex molecular assays that are difficult to implement in routine clinical laboratory practice. The authors of this study rationalized that because chromosomal structural changes and mutations ultimately converge on altered RNA expression patterns, an RNA-based classification should be both biologically comprehensive and technically more practical for clinical deployment.
This paper presents a novel strategy: rather than defining biological subgroups first and then assessing their clinical behavior, the authors reversed the process. They first grouped patients by their actual survival outcomes and then identified the gene expression biomarkers that could predict those groups. This survival-first approach was designed to maximize the clinical relevance of the final classification system.
The study enrolled patients from 22 medical centers participating in the DLBCL Consortium Program, which was approved by the institutional review board of each site and conducted in accordance with the Declaration of Helsinki. Informed consent was waived because of the retrospective study design. The primary cohort comprised 379 patients with de novo DLBCL; a completely separate cohort of 247 patients with extranodal DLBCL was reserved exclusively for external validation. All patients across both cohorts had been treated with R-CHOP. Patients with transformed DLBCL, primary mediastinal large B-cell lymphoma, or primary cutaneous DLBCL were excluded to maintain diagnostic homogeneity.
RNA library construction and sequencing: RNA was extracted from formalin-fixed paraffin-embedded (FFPE) tissue lysates using the Agencourt FormaPure Total 96-Prep Kit on an automated KingFisher Flex platform, enabling simultaneous DNA and RNA extraction from the same archived specimens. Libraries were selectively enriched for 1,408 cancer-associated genes using Illumina TruSight RNA Pan-Cancer Panel reagents. cDNA was generated from cleaved RNA fragments using random primers during first- and second-strand synthesis, followed by adapter ligation and sequence-specific probe capture. Sequencing was performed on an Illumina NextSeq 550, requiring 10 million reads per sample at 2x150 bp read length. Sequencing depth ranged from 10x to 1,739x with a median of 41x. Expression levels were quantified as fragments per kilobase of transcript per million (FPKM) using Cufflinks.
Patient demographics in the primary cohort: Of the 379 patients, 239 (63%) had nodal disease and 140 (37%) had extranodal disease. The International Prognostic Index (IPI) was greater than 2 in 141 patients (37%) and 2 or below in the remaining 238 (63%). Eastern Cooperative Oncology Group (ECOG) performance status was above 1 in only 60 patients (16%), with the majority (84%) performing well. Males constituted 55% of the cohort. All 379 patients were also classified by COO, providing a reference framework for subsequent correlation analyses.
The use of FFPE-derived material is a notable methodological strength, as FFPE represents the most commonly available tissue format in clinical practice. Demonstrating that targeted RNA sequencing from FFPE performs reliably enough to support machine learning-based classification directly addresses a key barrier to clinical translation.
The machine learning pipeline began not with gene expression data but with survival data alone. A fundamental obstacle in survival analysis is that censored cases - patients whose follow-up ended before the event of interest - do not have known survival times and therefore cannot be used directly as labeled training data for supervised classification. The authors addressed this by deriving expected survival times for censored patients from the Kaplan-Meier survival curve. Specifically, for a patient censored at time t0, the conditionally expected survival is estimated as t0 plus a weighted integral of the survival function beyond t0, normalized by S(t0). To avoid systematic bias toward overestimating survival (which the pure conditional expectation introduces), survival for cases censored before the median was set to the overall mean, while the full integral formula was applied to cases censored after the median.
Hierarchical two-step grouping: Using these estimated survival times, the 379 patients were first divided into two groups: short survival (S) and long survival (L), with a hazard ratio of 0.237 (95% CI: 0.170-0.330, P less than 0.00001). Inspection revealed that survival within each group was still heterogeneous, so each group was further split, generating four final groups: LL (long survival within the long group), LS (short survival within the long group), SL (long survival within the short group), and SS (short survival within the short group). The four-group model achieved a hazard ratio of 0.174 (95% CI: 0.120-0.251, P less than 0.0001), representing substantially improved discrimination.
Generalized naive Bayesian classifier: After defining the survival groups, the authors developed a generalized naive Bayesian classifier to identify which of the 1,408 measured genes could predict group membership. Standard naive Bayesian classifiers suffer from numerical underflow when feature dimensionality is high, because the likelihood product of many small probabilities approaches zero and loses information. The authors solved this by applying a geometric mean transformation (raising the likelihood to the power of 1/d, where d is the number of features), proven mathematically to prevent underflow while preserving the relative ordering of class probabilities. Cross-validation across 12 randomly split subsets of the 378-patient training set was used at each selection step to ensure biomarker robustness and minimize overfitting.
Feature selection across three steps: The gene selection process operated hierarchically across the three classification steps: 60 genes for S vs. L, 60 genes for LL vs. LS, and 60 genes for SL vs. SS, for a total of 180 genes. Critically, there was minimal overlap among the three gene sets, confirming that each step captured genuinely distinct biological signals. The full gene lists are detailed in Supplementary Table S1 of the paper.
Internal validation demonstrated that the 180 selected genes accurately reproduced the four predicted survival groups in the original 379-patient cohort. Kaplan-Meier curves for both overall survival (OS) and progression-free survival (PFS) showed clear, statistically significant separation across all four groups as expected by the machine learning assignment, confirming that the biomarker-based classification faithfully reproduced the survival groupings derived from clinical outcome data alone.
External validation in extranodal DLBCL: The critical test of the model was its performance in an entirely independent cohort of 247 extranodal DLBCL patients. When the selected biomarkers were applied to this cohort and patients were divided into two groups using the first-step (S vs. L) gene set, the model achieved a hazard ratio of 0.26 (95% CI: 0.278-0.653, P = 0.002). Applying all three gene sets to divide the extranodal cohort into four groups yielded a hazard ratio of 0.530 (95% CI: 0.234-1.197, P = 0.005). Although extranodal DLBCL characteristically carries a shorter overall survival than nodal disease - a finding confirmed in this cohort - the model maintained its ability to stratify patients within this group, demonstrating cross-population generalizability.
Combined cohort analysis: As a further robustness check, the two patient groups were merged into a combined cohort of 626 patients. Two-thirds were used to rebuild the model and one-third were held out as a test set. The overall model structure was preserved in the test set, continuing to identify two groups with intermediate survival but biologically distinct gene expression backgrounds, consistent with the LS and SL designations in the original four-group structure.
A particularly telling finding emerged from comparing the LS and SL groups. Despite having similar overall survival curves, these two groups required completely non-overlapping gene sets for their respective predictions. This genetic dissimilarity between two clinically similar patient groups underscores the substantial biological heterogeneity of DLBCL and argues against single-biomarker explanations for clinical behavior.
A series of multivariate Cox proportional hazard regression analyses were conducted to test whether the transcriptomic survival classification provided independent prognostic information beyond established clinical and molecular markers. In the first multivariate model incorporating survival classification, COO (GCB vs. ABC), and IPI, both the survival classification (HR 1.73, P less than 0.000001) and IPI (HR 2.41, P less than 0.000001) were independent predictors of survival, while COO was not (HR 0.97, P = 0.869). This finding indicates that the survival classification subsumes the prognostic information carried by COO subtype, providing a more powerful and inclusive risk stratification tool.
TP53 mutation: Of the 379 patients, 82 (22%) harbored TP53 mutations, which were significantly enriched in the short survival groups (P = 0.009). When TP53 mutation status was added to the multivariate model alongside survival classification, IPI, and COO, TP53 mutation remained an independent predictor of worse survival (HR 1.44, P = 0.048). Across multiple multivariate model configurations, TP53 mutation was the only molecular variable that retained independent prognostic significance after accounting for the survival classification. This positions TP53 as a uniquely important biomarker that captures biological risk not fully reflected in the transcriptomic survival groups.
MYD88 and CD79B mutations: MYD88 mutations were significantly more prevalent in the aggressive S group (P = 0.001), but when incorporated into a multivariate model with survival classification, IPI, COO, CD79B, and TP53, MYD88 mutation emerged as an independent predictor of better survival (HR 0.64, P = 0.042), suggesting that MYD88 mutation may partly identify patients who respond exceptionally well within an otherwise adverse survival group. CD79B mutation showed no prognostic significance in any model (HR 1.06, P = 0.842).
MYC and IRF4 expression: MYC mRNA was significantly elevated in the short survival groups (P less than 0.0001) and was higher in SL than in LS (P less than 0.0001), despite similar survival between these groups. However, MYC expression did not achieve independent predictive status in any multivariate model, whether treated as a continuous variable or categorized at the upper quartile threshold. Similarly, IRF4 mRNA was significantly overexpressed in the S group and was lower in LS than in SL, but reached only borderline independence in multivariate analysis (P = 0.067). This pattern suggests that MYC and IRF4 expression are partially co-linear with the transcriptomic survival classification and do not provide fully independent prognostic information beyond it.
Beyond their prognostic utility, the four survival groups carry biologically interpretable signals. The correlation with COO classification is instructive: the majority of GCB cases clustered in the favorable LL and LS groups (P less than 0.0001), while ABC cases predominated in the short survival groups - particularly SS. Importantly, although LS and SL groups showed overlapping overall survival distributions, the LS group contained significantly more GCB cases than the SL group (P = 0.016). This confirms that clinically similar outcomes in these two groups arise from distinct underlying tumor biologies, a conclusion reinforced by the non-overlapping gene sets used to classify each group.
The IRF4 finding: IRF4 gene translocation is classically associated with overexpression and with a recognized variant of large B-cell lymphoma that has been described as having a favorable clinical course in pediatric patients. In this adult DLBCL cohort, IRF4 mRNA overexpression was paradoxically more common in the short survival S group, and the SL subgroup showed significantly higher IRF4 levels than the LS subgroup (P = 0.02). This apparent contradiction with some prior literature may reflect the distinct clinical context - adult de novo DLBCL treated with R-CHOP - compared to the pediatric IRF4-rearranged lymphoma setting in which the favorable prognosis was originally described.
Age confounding in the SS group: The SS (worst survival) subgroup contained a significantly higher proportion of patients above age 60 (P = 0.01). The authors flag this as an important caveat: death from causes other than lymphoma in older patients may be contributing to the inferior survival designation of SS, potentially conflating poor lymphoma biology with competing mortality. This observation does not invalidate the survival classification but underscores that prognostic models in older populations should ideally account for competing risks.
The identification of specific gene sets driving each survival subgroup (listed in Supplementary Table S1) also provides a foundation for hypothesis-driven targeted therapy development. Each survival-defined subgroup is associated with a distinct set of biologically active genes, offering potential drug targets that could be matched to patients in clinical trials on the basis of their survival-group membership and the underlying pathway activity reflected in their RNA profile.
The practical goal of this classification system is to identify patients with DLBCL who are unlikely to respond to R-CHOP before they receive it, so they can be enrolled in clinical trials or receive alternative upfront therapy. The current standard approach - treating all R-CHOP-eligible patients with the same regimen and reassessing at interim PET - means non-responders are exposed to toxic therapy that will not cure them. A reliable upfront predictor of non-response would fundamentally alter the treatment algorithm.
Automation and clinical deployability: The authors explicitly note that their subclassification pipeline has been implemented as software that, when fed RNA sequencing data from targeted transcriptomics, automatically outputs a survival group assignment. This is a critical distinction from prior molecular classifiers that require manual interpretation or complex computational infrastructure. The use of targeted sequencing from FFPE tissue - the standard pathological format - rather than fresh-frozen specimens or whole-exome sequencing also meaningfully lowers the barrier to clinical implementation. Targeted RNA-seq panels are already available on clinical-grade sequencing platforms.
Comparison with existing tools: The IPI remains an independent predictor in this model but captures only clinical variables (age, stage, LDH, performance status, extranodal sites). The transcriptomic survival classification adds a layer of molecular resolution that IPI cannot provide. The fact that COO classification lost prognostic independence once the survival classification was included suggests that the 180-gene panel effectively integrates and supersedes the information captured by GCB/ABC subtyping alone, while also encoding additional prognostic signal from other pathways.
Enriching clinical trials with patients from specific survival subgroups is another application the authors highlight. If patients with similar biology and clinical trajectory are grouped together in trials, the biological signal of a new therapy is more likely to emerge against a homogeneous background, potentially improving trial sensitivity and reducing sample size requirements for demonstrating benefit.
Retrospective design and cohort composition: Both the training and validation cohorts are retrospective, drawn from patients treated at academic and specialized centers across the DLBCL Consortium Program. Patients with transformed, mediastinal, or cutaneous DLBCL were excluded, which limits direct applicability to these clinically important subtypes. The extranodal validation cohort confirmed cross-cohort generalizability, but prospective validation in a community oncology setting - where most DLBCL patients receive care - has not yet been performed.
Age confounding in the SS group: As noted in the multivariate analysis, the poorest survival group (SS) is enriched with patients above age 60. Since this is a real-world R-CHOP-treated population, some deaths in this group likely reflect non-lymphoma causes rather than disease-specific mortality. The current survival model does not formally account for competing risks, meaning the SS designation may partially reflect comorbidity burden or treatment tolerance rather than pure tumor biology. Competing-risk models or cause-specific survival analyses would strengthen future iterations.
TP53 as a residual independent signal: The persistence of TP53 mutation as an independent adverse prognostic factor even after controlling for the transcriptomic survival classification is both practically important and biologically interesting. It suggests that TP53 mutation contributes to poor outcome through mechanisms not fully captured by the RNA expression landscape of 1,408 genes - possibly through genomic instability, altered DNA damage response, or therapy resistance pathways that do not robustly manifest as steady-state RNA changes. Integrating TP53 mutation status directly into a combined genomic-transcriptomic classifier is a natural next step.
Future directions: The authors point toward prospective clinical trials in which patients are stratified by survival subgroup at diagnosis, with non-favorable groups (SL, SS) enrolled in intensified or novel therapy arms. Expanding the targeted panel to capture additional potentially prognostic genes, integrating mutation profiling with RNA expression in a single assay, and exploring the gene lists associated with each survival subgroup as actionable therapeutic targets are all areas of active development. Validating the pipeline across ethnically and geographically diverse patient populations beyond the current 22-center consortium will be essential before regulatory submission.