Artificial Intelligence and Classification of Mature Lymphoid Neoplasms

Exploration of Targeted Anti-tumor Therapy 2024 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Why Lymphoma Classification Is Complex and Where AI Enters

Mature lymphoid neoplasms encompass a broad spectrum of B-cell, T-cell, and NK-cell malignancies that are among the most diagnostically challenging tumors in all of oncology. Unlike many solid tumors that can be classified primarily by anatomical location and histological grade, lymphoid neoplasms require integrating morphological features, immunophenotypic profiles, cytogenetic data, and molecular pathology findings into a unified diagnosis. This multidimensional requirement has driven three decades of iterative international classification efforts, each revision expanding the number of recognized entities as new molecular markers reveal biologically and clinically distinct subgroups.

Classification timeline: The modern framework traces back to the 1994 Revised European-American Classification of Lymphoid Neoplasms (REAL), which was subsequently reorganized under the World Health Organization (WHO) umbrella in 2001, updated in 2008, and revised again in 2016. The 2016 revision is the classification that most practicing hematopathologists currently use. In 2022, two parallel updates emerged: the International Consensus Classification (ICC) of mature lymphoid neoplasms, issued by the Clinical Advisory Committee, and the 5th edition of the WHO Classification of Haematolymphoid Tumours (WHO-HAEM5). Because both documents derive from the same 2016 foundation, they share most major categories but differ in diagnostic criteria for specific entities and in how they handle several emerging or provisional subtypes.

Role of AI: This paper, authored by Carreras, Hamoudi, and Nakamura from Tokai University School of Medicine and the University of Sharjah, serves two functions. First, it provides a comparative review of the ICC 2022 and WHO-HAEM5 classifications for mature B-cell neoplasms, highlighting the key differences. Second, it summarizes the authors' body of AI research using machine learning and neural networks to classify lymphoma subtypes and predict patient prognosis. The paper culminates in a new analysis using a multilayer perceptron (MLP) neural network to classify B-cell lymphoma subtypes from a panel of conventional cell-of-origin markers currently used in routine clinical immunohistochemistry.

The paper is framed as both a clinical resource (for understanding where the two major classifications agree and diverge) and a proof-of-concept demonstration (showing that AI can leverage conventional diagnostic markers for automated lymphoma subtyping). Published in Cancers in 2024, it draws on more than a decade of the authors' accumulated AI research in hematopathology.

TL;DR: This 2024 paper reviews the ICC 2022 and WHO-HAEM5 lymphoma classifications, compares their key differences for mature B-cell neoplasms, and summarizes AI research using machine learning and neural networks to classify lymphoma subtypes and predict prognosis. A new MLP analysis achieves 79% overall classification accuracy using 12 routine immunohistochemistry cell-of-origin markers.
Pages 2-5
ICC 2022 vs. WHO-HAEM5: Where the Two Classifications Agree and Diverge

The paper provides a detailed comparison table (Table 1) mapping corresponding entities across the ICC 2022, WHO-HAEM5, and the prior WHO revised 4th edition. The most "relevant subtypes" recognized by all three frameworks include chronic lymphocytic leukemia/small lymphocytic lymphoma (CLL/SLL), splenic marginal zone lymphoma (MZL), hairy cell leukemia, lymphoplasmacytic lymphoma (LPL) and Waldenstrom macroglobulinemia, multiple myeloma, extranodal MZL of mucosa-associated lymphoid tissue (MALT lymphoma), nodal MZL, follicular lymphoma (FL), mantle cell lymphoma (MCL), diffuse large B-cell lymphoma (DLBCL), and Burkitt lymphoma (BL), among others. These principal subtypes are well-established and consistently represented across all three classification systems.

Key new and modified entities: The ICC introduced several notable changes marked with an asterisk. Hairy cell leukemia-variant is renamed to "splenic BCL/leukemia with prominent nucleoli," acknowledging that some of these cases overlap with B-cell prolymphocytic leukemia. Primary cutaneous marginal zone lymphoproliferative disorder is elevated to a distinct entity, previously considered a subtype of extranodal MZL of MALT. HBCL with MYC and BCL2 rearrangements, and separately HBCL with MYC and BCL6 rearrangements, are now formally separated. WHO-HAEM5 treats BCL6 rearrangement as less relevant and reclassifies MYC/BCL6 cases as either DLBCL NOS or HGBL NOS based on cytomorphology, creating a meaningful divergence from the ICC approach.

Follicular lymphoma grading: One of the most practically significant differences between the two classifications concerns follicular lymphoma grading. The ICC 2022 retains morphological grading (grades 1-2, 3A, and 3B), while WHO-HAEM5 makes grading no longer mandatory in classic FL cases. WHO-HAEM5 instead defines two new FL subtypes: follicular large B-cell lymphoma (FLBL, formerly FL grade 3B) and FL with uncommon features (uFL). This divergence matters clinically because grade 3B FL has historically been treated with the same regimens used for DLBCL, so reclassifying it as a separate entity rather than a grading designation could affect treatment guidelines.

B-cell differentiation biology: The paper situates the major B-cell lymphoma subtypes within normal B-cell development stages, a framework that underpins both classification and AI approaches. MCL arises from pre-germinal center mantle zone B cells. FL, BL, and DLBCL develop from germinal center B cells, with the hallmark FL translocation t(14;18)(q32;q21) occurring in bone marrow before cells enter the germinal center. MALT lymphoma derives from post-germinal center marginal zone B cells, and plasma cell myeloma from long-lived post-germinal center plasma cells. Understanding these developmental origins is directly relevant to the cell-of-origin markers used in the AI classification analysis later in the paper.

TL;DR: ICC 2022 and WHO-HAEM5 share most major lymphoma categories but diverge on several key points: FL grading is retained in ICC but made optional in WHO-HAEM5 (which adds FLBL and uFL subtypes), and MYC/BCL6 rearrangement cases are handled differently. New entities in ICC include primary cutaneous MZL as a distinct subtype and formal separation of HBCL with MYC/BCL2 vs. MYC/BCL6 rearrangements.
Pages 5-6
Molecular Subtyping of DLBCL: Genetic Classification Beyond Cell-of-Origin

DLBCL NOS is the most common aggressive lymphoma and remains one of the most heterogeneous entities in the classification. While cell-of-origin (COO) subtyping into germinal center B-cell (GCB) and activated B-cell (ABC) subtypes has been the cornerstone of DLBCL molecular classification for two decades, the ICC 2022 recognizes that more granular molecular profiling using 5 or 7 functional subgroups can refine prognostic and potentially therapeutic stratification. This section of the paper reviews the two most influential genetic classification systems for DLBCL.

Chapuy-Shipp classification (C0-C5): The Chapuy-Shipp system identified five molecular clusters in 304 DLBCL samples based on coordinated genetic signatures. Cluster C0 is characterized by an absence of molecular changes. C1 by BCL6 structural variants. C2 by TP53 mutations, associated with intermediate prognosis. C3 by BCL2 mutations and structural variants creating IGH/BCL2 juxtaposition, along with frequent KMT2D, CREBBP, and EZH2 mutations, with a GCB-like cell-of-origin. C5 by 18q and chromosome 3 gains, and CD79B, MYD88, and PIM1 mutations with an ABC-like cell-of-origin. Progression-free survival differed significantly between favorable (C0/C1/C4), intermediate (C2), and poor (C3/C5) groups.

Wright classification (7 subtypes): The Wright probabilistic algorithm classifies DLBCL into seven genetic subtypes, MCD, N1, A53, BN2, ST2, EZB, and MYC+/MYC-, each with distinct molecular drivers and potential therapeutic targets. The MCD subtype is defined by MYD88 (L265P) and CD79B mutations, making it a candidate for BTK inhibitor therapy. BN2 by BCL6 fusions and NOTCH2 mutations, with a favorable prognosis. EZB by BCL2 fusions and EZH2/TNFRSF14 mutations, suggesting sensitivity to BCL2 and EZH2 inhibitors. ST2 by TET2 mutations. A53 by TP53 mutations. N1 by NOTCH1 mutations. This rationally targeted classification approach directly informed the design of clinical trials testing BTK, PI3K, BCL2, JAK, IRF4, and EZH2 inhibitors in molecularly selected DLBCL patients.

The ICC 2022 maintains the COO sub-classification (GCB and ABC) for routine clinical use, while acknowledging that the molecular profiling schemas represent a more sophisticated layer of stratification relevant to clinical trial design and emerging targeted therapies. For most clinical centers, the 5 or 7 functional subgroups are research-grade tools rather than routine diagnostic tests, but this is expected to change as next-generation sequencing becomes more accessible.

TL;DR: DLBCL molecular subtyping goes beyond GCB/ABC COO into 5 Chapuy-Shipp clusters (C0-C5) and 7 Wright subtypes (MCD, BN2, EZB, ST2, A53, N1, MYC+/-). MCD responds to BTK inhibitors, EZB to BCL2/EZH2 inhibitors. BN2 carries the most favorable prognosis. These classifications guide targeted therapy selection and clinical trial design.
Pages 6-8
Machine Learning Algorithms and Neural Network Architectures Used in Lymphoma AI Research

The paper provides a concise taxonomy of the AI techniques the authors employed across their series of lymphoma research publications. The distinction between "strong AI" (artificial general intelligence, AGI) and "weak AI" (narrow AI) frames the clinical AI landscape: strong AI remains theoretical and would need to pass the Turing test, while weak AI focuses on specific, well-defined tasks. All practical medical AI applications, including every technique described in this paper, are forms of weak AI.

Machine learning methods used: The authors' research portfolio applied a wide range of machine learning techniques, including the C5 decision tree (robust to missing data, handles large numbers of predictors, short training time), Bayesian networks (graphical models linking variables via nodes and arcs, resilient to missing data), C and R trees (handles missing data and continuous or categorical targets), CHAID decision trees (chi-squared automatic interaction detection, supports non-binary splits), discriminant analysis, k-Nearest Neighbor (KNN), logistic regression, linear support vector machines (LSVM), QUEST trees (binary classification with faster processing than C and R trees), random forest (bagging algorithm, highlights most relevant predictors), random trees, SVM, tree-AS, and XGBoost (extreme gradient boosting) in both linear and tree variants. This comprehensive list reflects the authors' strategy of applying multiple algorithms to the same datasets and comparing performance rather than committing to a single method.

Artificial neural networks: Two neural network architectures were used: multilayer perceptron (MLP) and radial basis functions. The MLP is a feed-forward, supervised learning model. Each neuron receives multiple inputs, multiplies each by a weight, adds a bias term, and passes the result through a nonlinear activation function (such as the sigmoid function). A minimum architecture requires an input layer of predictors, one or more hidden layers with unobservable nodes, and an output layer with the predicted variable. The network learns by adjusting weights during training to minimize prediction error. In the context of lymphoma subtyping, the input layer consisted of gene expression values and the output layer predicted categorical lymphoma subtype labels.

Dataset: The lymphoma classification analysis used the publicly available gene expression dataset GSE132929. Predictors were selected specifically to mirror the immunohistochemical markers that hematopathologists already use in routine clinical diagnosis, including genes corresponding to CD3, CD5, CD19, CD79A, MS4A1 (CD20), MME (CD10/CALLA), BCL6, IRF4 (MUM-1), BCL2, SOX11, MNDA, and FCRL4 (IRTA1). This deliberate choice of clinically familiar markers was designed to make the AI model interpretable and directly relevant to practicing pathologists.

TL;DR: The authors applied 14+ machine learning methods (C5 tree, Bayesian network, random forest, XGBoost, SVM, KNN, etc.) and two neural network architectures (MLP and radial basis functions) across their lymphoma AI research series. Dataset GSE132929 was used for the new classification analysis, with 12 input variables matching routine IHC markers (CD3, CD5, CD19, CD20, CD10, BCL6, IRF4, BCL2, SOX11, MNDA, and others).
Pages 8-10
How AI Models Predicted Lymphoma Survival and Uncovered Prognostic Genes

The paper summarizes nine key publications from the authors' AI research program spanning 2020 to 2022 (Table 2). These studies applied machine learning and neural networks to gene expression datasets to predict overall survival and classify subtypes across the most frequent non-Hodgkin lymphoma entities, including DLBCL, FL, MCL, BL, and MZL. The approach was systematic: use high-dimensional gene expression data as predictors, train multiple AI models to identify the most informative features, then validate the candidate genes or proteins using independent techniques such as immunohistochemistry and gene-set enrichment analysis (GSEA).

DLBCL gene discovery pipeline: The foundational 2020 publication stratified 100 DLBCL cases by CD163 expression (a marker of M2-like tumor-associated macrophages). Using MLP with 54,614 gene probes as input predictors and overall survival (dead vs. alive) as the output, 25 prognostically relevant genes were identified. Correlation analysis, GSEA, and Cox regression then refined this to three key genes: MYC (cell cycle), BCL2 (apoptosis), and ENO3 (cell metabolism). High co-expression of MYC, BCL2, and ENO3 was associated with poor prognosis. This three-gene signature was subsequently validated by immunohistochemistry in independent DLBCL cases.

Dimensionality reduction and multivariable prediction: A follow-up study improved on this pipeline by predicting not just overall survival but also multiple clinicopathological variables simultaneously, including COO molecular subtype, age, LDH ratio, ECOG performance status, clinical stage, and extranodal disease. This multi-output approach refined the gene list further, ultimately highlighting 16 prognostic genes. From this expanded list, PD-L1 (CD274) and IKAROS emerged as the most clinically relevant and were validated by immunohistochemistry in an independent series from Tokai University Hospital, confirming their protein-level prognostic significance in DLBCL.

Pan-cancer and other lymphoma subtypes: The research extended beyond DLBCL. For MCL, two strategies were applied: first, dimensionality reduction based on previously identified prognostic genes; second, an immuno-oncology panel approach. MCL survival results were correlated with the Lymphoma/Leukemia Molecular Profiling Project (LLMPP) MCL35 proliferation assay, providing external benchmark validation. For FL, 120 independent artificial neural networks were trained using random number generator-based dimensionality reduction. A pan-cancer survival analysis was also performed across multiple tumor types using a cancer transcriptome panel, demonstrating the generalizability of the neural network approach beyond lymphoma.

TL;DR: From 54,614 gene probes across 100 DLBCL cases, MLP identified 25 prognostic genes, refined to MYC, BCL2, and ENO3 as the key survival-associated trio. A subsequent multi-output model highlighted PD-L1 and IKAROS, both validated by IHC in independent cases. MCL results were benchmarked against the LLMPP MCL35 assay. Studies extended to FL (120 neural networks), BL, MZL, and a pan-cancer transcriptome panel.
Pages 10-12
Neural Networks Classify Lymphoma Subtypes Without Histological Images

One of the most striking findings in the authors' research program is that neural networks trained exclusively on gene expression data, without any histological images, can classify multiple lymphoma subtypes with high accuracy. In their 2021 study, the input layer contained the gene expression values of a cancer transcriptome panel of 1,769 genes, and the output layer predicted one of five lymphoma subtype labels: FL, MCL, DLBCL, BL, and MZL. Using a neural network trained on all array genes or the pan-cancer panel, the model successfully classified all five subtypes with high overall performance. This demonstrated that the molecular signatures encoded in transcriptomic data are sufficient to replicate the diagnostic distinctions that pathologists make using morphological and immunophenotypic criteria.

Bayesian network for prognosis: A Bayesian network approach was applied to DLBCL prognosis prediction, providing a graphical probabilistic model of the relationships between gene expression nodes and survival outcome. The Bayesian network is particularly useful when data are missing or incomplete, since it generates the best feasible prediction from whatever input is available. The visual representation of the model (showing nodes for both predictors and targets connected by arcs) offers a degree of interpretability that pure deep learning models lack, making it more accessible for clinicians who want to understand the causal relationships the model has learned.

Immunohistochemistry and machine learning integration: The CASP8 study extended the AI approach to protein-level data obtained from digital image quantification of immunohistochemical stainings in DLBCL. Machine learning and neural networks were used to evaluate the prognostic value of CASP8 and 12 related markers (CASP3, cleaved PARP1, BCL2, TP53, MDM2, MYC, Ki67, E2F1, CDK6, MYB, LMO2, and TNFAIP8). High expression of CASP8 was found to correlate with favorable prognosis, a finding that would not have been identified through standard univariable analysis given the complexity of the marker interactions. This study demonstrates that AI can extract prognostic information from routine diagnostic IHC panels that standard statistical methods would miss.

Extension to non-tumor contexts: The same methodology was successfully applied to non-tumor immunological conditions, including celiac disease and ulcerative colitis, using autoimmune discovery transcriptomic panels. This cross-disease applicability validates the robustness of the approach and suggests that the AI pipeline is not overfit to lymphoma-specific biology, but instead captures general patterns in high-dimensional transcriptomic data relevant to a range of immunological disorders.

TL;DR: Neural networks using a 1,769-gene cancer transcriptome panel classified FL, MCL, DLBCL, BL, and MZL subtypes without histological images. Bayesian networks provided interpretable probabilistic prognosis models for DLBCL. Machine learning applied to CASP8 and 12 related IHC markers identified high CASP8 as a favorable prognostic marker. The same pipeline was validated in celiac disease and ulcerative colitis.
Pages 12-14
Classifying Lymphoma Subtypes from Routine Immunohistochemical Markers Using a Neural Network

The central new contribution of this paper is a neural network analysis that classifies mature B-cell lymphoma subtypes using only the 12 cell-of-origin marker genes that hematopathologists routinely test by immunohistochemistry. These markers were selected because they reflect the postulated cellular origin and differentiation stage of each lymphoma subtype: CD5 and CD3E (T-cell markers that help exclude T-cell lymphomas), BCL2 (anti-apoptotic, germinal center and post-germinal center marker), BCL6 (germinal center transcription factor), IRF4/MUM-1 (post-germinal center plasma cell differentiation marker), MME/CD10/CALLA (germinal center marker), CD19 and MS4A1/CD20 (B-lymphocyte markers), CD79A (B lymphocyte antigen receptor complex), SOX11 (transcriptional activator, highly specific for MCL), MNDA (myelomonocytic and marginal zone B-cell marker), and FCRL4/IRTA1 (memory B-cell function marker).

Neural network architecture and results: The MLP model was trained using the publicly available dataset GSE132929. The input layer contained 12 nodes (one per marker gene), and the number of hidden layers was determined automatically during training. The output layer predicted lymphoma subtype across five categories: FL, MCL, DLBCL, BL, and MZL. The overall correct classification rate was 79%. Broken down by subtype, MCL was the best predicted at 88%, followed by FL at 85%, BL at 80%, DLBCL at 79%, and MZL at 44%. The classification matrix (Figure 3) shows the distribution of correct and incorrect predictions for each subtype.

Why MZL is hardest to classify: The low accuracy for MZL (44%) is diagnostically meaningful and not simply a modeling failure. MZL is recognized as one of the most challenging lymphoma diagnoses in routine pathology practice. The same IHC markers used as inputs to the neural network are the same markers histopathologists use clinically, and in borderline or overlap cases, pathologists themselves often render a diagnosis of "indolent mature BCL, unspecified." The fact that the neural network struggles with MZL in the same way reflects genuine biological overlap between MZL and other indolent B-cell lymphomas rather than a limitation of the AI approach per se.

Clinical implication: This analysis demonstrates that a simple 12-marker gene expression panel, directly mirroring routine clinical IHC practice, can achieve 79% overall lymphoma subtype classification. When combined with histological assessment and cytogenetics, such a model could serve as a computational second-opinion tool or as a standardized quantitative input for AI-assisted diagnostic pipelines. The transparency of the input markers, all well-established in clinical pathology, addresses a key barrier to clinical adoption that has hampered black-box deep learning models.

TL;DR: A 12-input MLP using routine IHC cell-of-origin markers (CD5, CD3, BCL2, BCL6, IRF4, CD10, CD19, CD20, CD79A, SOX11, MNDA, FCRL4) achieved 79% overall lymphoma subtype accuracy on dataset GSE132929. MCL was classified best at 88%, FL at 85%, BL at 80%, DLBCL at 79%, and MZL at 44%. Low MZL accuracy mirrors the recognized diagnostic difficulty in clinical practice.
Pages 14-17
AI as a Future Bioinformatics Tool in Lymphoma Classification

The paper's conclusions draw together two parallel threads: the state of lymphoma classification and the state of AI in hematopathology. On the classification side, the authors note that while ICC 2022 and WHO-HAEM5 are largely concordant, reflecting their common origin in the 2016 WHO classification, there are conceptual differences and differences in diagnostic criteria for specific entities that practicing pathologists need to navigate. The trend across all classification revisions is toward incorporating more molecular data, from immunophenotype and cytogenetics toward full genomic profiling, as next-generation sequencing becomes a standard diagnostic tool.

AI as a bioinformatics layer: The authors argue that AI should be viewed not as a replacement for pathologist expertise but as a bioinformatics tool, sitting alongside GSEA, Cox regression, and transcriptomic profiling in the analytical toolkit. Neural networks excel at pattern recognition across high-dimensional datasets, identifying non-obvious combinations of features that predict outcome or subtype more accurately than any single marker. The research summarized in this paper demonstrates that these capabilities apply directly to lymphoma, where the biological complexity exceeds what conventional statistical approaches can fully capture.

Limitations of the current work: The new cell-of-origin analysis uses a single publicly available dataset (GSE132929) and reports only training/internal performance without external validation. The 44% accuracy for MZL classification highlights that a 12-marker panel, while informative, is insufficient for all lymphoma subtypes. Additionally, the analysis works with gene expression data (RNA-level) rather than protein-level IHC data, so the relationship to actual clinical workflow requires further validation. The restriction to five lymphoma categories (FL, MCL, DLBCL, BL, MZL) leaves many WHO-recognized entities unaddressed.

Trajectory: The authors envision that AI will be incorporated into future lymphoma classification frameworks as a bioinformatics component alongside genomic profiling, in a manner analogous to how molecular testing was gradually integrated into the WHO classification over four successive editions. This integration could take the form of AI algorithms that generate probabilistic subtype assignments from gene expression or mutational profiles, supplementing rather than replacing the expert pathologist review that remains central to lymphoma diagnosis. As genomic profiling moves from research into routine clinical practice, the datasets needed to train and validate such models will become progressively more accessible.

TL;DR: ICC 2022 and WHO-HAEM5 are largely concordant but differ on FL grading, MYC/BCL6 rearrangement handling, and several emerging entities. AI is positioned as a bioinformatics tool complementing pathologist expertise, not replacing it. Current limitations include single-dataset validation, only five lymphoma categories covered, and gene expression rather than protein-level data. AI integration into classification frameworks is projected to follow the same incremental path as molecular testing adoption across four WHO classification editions.