Digital pathology (DP) is the process of converting histology glass slides to digital images using sophisticated computerized technology, transforming analog histologic data into a format amenable to analysis by machine learning (ML) and artificial intelligence (AI). This 2020 review from the University of Texas MD Anderson Cancer Center, published in Cancers, surveys the state of these emerging technologies in the context of diagnostic hematology and the evaluation of lymphoproliferative diseases, a field where AI applications have historically lagged behind oncology disciplines such as dermatology, radiology, and prostate pathology.
Machine learning framework landscape: The authors categorize ML approaches into traditional methods, including support vector machines (SVM), neural networks (NN), logistic regression, random forest, and naive Bayes classifiers, and modern deep learning (DL) methods built around convolutional neural networks (CNN or ConvNet), long short-term memory (LSTM) networks, deep belief networks, and encoder-decoder architectures. Deep learning algorithms emulate the layered structure of human brain neurons, with each group of neurons receiving input from the layer below, and learning to detect characteristic morphological features through supervised, unsupervised, or semi-supervised training. CNNs specifically excel because they include convolution layers for hierarchical feature extraction and pooling layers for dimensionality reduction, allowing the network to deal efficiently with the enormous feature spaces inherent in high-resolution microscopy images.
Regulatory milestones: A key backdrop to this review is the FDA clearance of the Philips IntelliSite Pathology Solution as the first whole-slide imaging (WSI) system approved for primary clinical diagnostic evaluation in the United States. Subsequently, Leica Biosystems received clearance to market the Aperio AT2 DX System for clinical diagnosis. These approvals marked an inflection point for computational pathology, pushing the boundaries of what is possible in terms of AI integration into routine clinical workflows. The CellaVision DM96 automatic hematology analyzer, introduced in 2004 and succeeded by the DM1200 (2009) and DM9600 (2014), represents an earlier generation of FDA-cleared digital tools already operating in clinical hematology laboratories.
Despite these advances, the review notes that studies of DP applications specifically in hematopathology remain limited relative to other fields. Contributing factors include the need for high-power magnification and the historically perceived reliance on oil-immersion objectives, as well as the inherent complexity of hematologic diagnoses that rely on integrating immunophenotypic, cytogenetic, flow cytometric, and morphological data.
Accurate identification of peripheral blood elements and leukocyte subsets is a fundamental component of hematologic workup. Manual analysis of peripheral blood smears (PBS) is time-consuming, labor-intensive, and subject to intra- and interobserver variability. Automated analyzers using digital image analysis have made significant inroads, with the CellaVision software grouping red blood cells (RBCs) into 21 morphological categories. Particularly consequential is the schistocyte group, given that schistocyte detection raises concern for microangiopathic hemolytic anemia and thrombotic thrombocytopenia purpura (TTP). Studies comparing schistocyte counts between CellaVision DM96 and conventional manual microscopy found the automated system achieves high sensitivity but poor specificity, requiring reclassification by an expert laboratory professional. This low specificity is attributed to the broad and poorly defined morphological spectrum of schistocytes, a problem that AI-based algorithms with dynamic learning capacity are well-positioned to address.
Lymphoid cell classification: Alferez et al. developed methods using Watershed Transformation image processing to automatically classify normal and mature B-cell neoplasms from peripheral blood, including chronic lymphocytic leukemia (CLL) and hairy cell leukemia (HCL). An initial study using 44 extracted features demonstrated high precision in distinguishing CLL from HCL cells. A subsequent 2015 expansion incorporating 113 extracted features combined with color image segmentation, tested on specimens from healthy individuals and patients with CLL, HCL, and mantle cell lymphoma (MCL), achieved an accuracy of 98.07%, with precision of 99.7%, sensitivity of 97.5%, and specificity of 98.6%.
Broadening the classification range: Adding color and texture features to geometric cytologic characteristics via a support vector algorithm, investigators classified three categories (reactive lymphocytes, normal lymphocytes, and abnormal lymphoid cells) with 97.67% accuracy. Further sub-classification of abnormal lymphocytes into HCL, MCL, FL, CLL, and prolymphocytic subtypes reduced accuracy to 91.23%, reflecting the increasing difficulty of finer-grained morphological distinctions. A more ambitious proof-of-concept incorporated 27 geometry features and 2,649 color and texture features to classify a wide array including B-PLL, SMZL, T-PLL, T-LGL, and Sezary syndrome cells, successfully defining entity-specific feature sets. For blast classification into myeloblasts or lymphoblasts, SVM-based methods achieved true positive rates of 85%, 82%, and 74% for reactive lymphoid cells, myeloblasts, and lymphoblasts respectively.
CNN for white blood cell detection: Convolutional neural networks trained to classify four WBC types (eosinophils, neutrophils, lymphocytes, monocytes) achieved 93% precision for mononuclear vs. polynuclear discrimination, dropping to 88% across all four classes. For real-time detection requirements, Wang and colleagues compared Single Shot Multibox Detector (SSD, 300x300 input) and YOLOv3 pipelines across 11 WBC categories. SSD achieved a mean average precision (mAP) of 93.1% vs. YOLOv3's 92.25%. Detection accuracy was 97% for blasts, nearly 100% for mature WBCs, and 87% for immature types. YOLOv3 demonstrated a major speed advantage with 14 ms inference time per image vs. 53 ms for SSD, highlighting the tradeoff between accuracy and computational speed for real-time applications.
The classification of leukemic diseases depends on integrating immunophenotypic, flow cytometric, cytogenetic, and mutational analysis results alongside histologic and cytologic features. Efforts are underway to develop tools for maximizing automated extraction of detailed information from peripheral blood and bone marrow (BM) smears as additional ancillary diagnostic tools at baseline and for disease monitoring. A key technical limitation of WSI in hematopathology has been the unavailability of three-dimensional images. Researchers from Stanford University addressed this using a super-resolution imaging pipeline, constructing a three-dimensional digital picture from multiple overlapping images of BM aspirate smears, achieving a significant improvement in sharpness and resolution.
Bone marrow cellularity assessment: Current evaluation of BM cellularity is largely subjective, reported based on visual estimates correlated with patient age. Two digital image analysis platforms have been evaluated for standardizing this measurement. The HALO imaging software, which reports multiplexed morphological data on a cell-by-cell basis, was assessed for BM trephine biopsy cellularity, yielding only 81% correlation (ICC = 0.81) with visual estimates by expert hematopathologists. The VisioPharm platform demonstrated stronger performance when applied to quantitation of lymphocytic aggregates in Sjogren biopsies, where a Bayesian-based algorithm matched pathologists' scoring in 100% of cases. These results highlight that the precision of image analysis tools is highly task-dependent.
Differential cell counts: BM differential cell counts (DCC) are a critical classification criterion for hematologic malignancies, but manual DCC is time-consuming and subject to major inter- and intra-observer bias. ML techniques for automated DCC on BM aspirates are in early development, with promising results emerging for identification and classification of normal BM elements. Enumeration of blasts, including distinguishing different blast subtypes, is pivotal for disease diagnosis and classification of acute leukemia. A systematic review of leukemia ML studies published after 2010 encompassed the four main types: acute lymphocytic leukemia (ALL), CLL, acute myeloid leukemia (AML), and chronic myeloid leukemia (CML), noting that most researchers used supervised DL algorithms, with a more recent shift toward unsupervised methods to reduce the manual segmentation burden.
Flow cytometry automation: Models for automated interpretation of flow cytometry (FC) results have been developed for diagnosing hematologic malignancies. Biehl et al. applied generalized matrix relevance learning vector quantization (GMLVQ) in 179 cases to achieve 100% accuracy for AML detection. Manninen et al. used a regularized logistic regression model (LDA) in 359 cases, also achieving 100%. Dundar et al. applied non-parametric Bayesian algorithms (ASPIRE) for AML classification with a 99% area under the curve (AUC). A separate group achieved 99.6% accuracy in FC diagnosis of CLL using Bayesian clustering on a dataset augmented to 50,000 virtual cases through resampling. SVM applied to CML immunophenotypic patterns identified malignant myeloid cells vs. normal/reactive neutrophils with high discrimination.
A substantial body of work has focused on AI-based diagnosis of acute lymphocytic leukemia (ALL), particularly subtyping on morphological grounds. Using supervised ML models on BM specimens only, Rehman et al. achieved an overall accuracy of 98% and Reta et al. achieved 92% for sub-classification into FAB subtypes L1, L2, and L3 vs. normal marrow. On the PBS side, Shafik et al. applied unsupervised deep CNN models and demonstrated 100% sensitivity and 98% specificity for ALL detection, and 97% sensitivity with 99% specificity for ALL sub-classification into L1, L2, and L3 vs. normal. Across studies using SVM-based supervised methods for lymphoblast detection and differentiation from reactive lymphoid cells, overall accuracy ranged from 74% to 99%, with maximum sensitivity reaching 100% and specificity up to 95%.
Multi-classifier comparison for ALL: The best outcomes were reported by Bhattacharjee et al., who collected 120 cases and applied pattern recognition-based segmentation to train and compare multiple classifiers, including artificial neural network (ANN), k-nearest neighbor (kNN), k-means, and SVM. Combined BM and PBS analysis achieved results comparable to PBS alone for overall leukemia detection, with further sub-classification showing sensitivity, specificity, and accuracy above 90% for distinguishing L1, L2, and L3 ALL subtypes from non-neoplastic cells. An AI study in ALL also encompassed therapy optimization using the Phenotypic Personalized Medicine Digital Health Platform, which identifies patient-specific factors correlating drug dosage with phenotypic outputs; the algorithm demonstrated that adjusted dosing of combination chemotherapy could enhance outcomes while reducing the quantity of chemotherapy administered.
AML detection and subtyping: Automated detection and sub-categorization of AML has also been investigated. Pattern recognition-based supervised algorithms and SVM classifiers have achieved myeloblast detection accuracy ranging from 82% to 97%. Reta et al. developed an algorithm to distinguish FAB subtypes M2, M3, and M5, reporting 100% accuracy for these three subtypes. However, the authors note that the newer WHO classification of AML incorporates numerous genetic-based entities, limiting the direct clinical applicability of FAB-based AI systems. The primary value of these preliminary studies lies in identifying the most efficient segmentation methods and classifiers for eventual integration into routine workflows.
CLL and atypical CLL identification: A 2014 case report from Memorial Sloan Kettering demonstrated that CellaVision digital microscopy could serve as a rapid screening tool for atypical CLL (aCLL), identifying 60% lymphocytes and 26% large abnormal lymphoid cells in a patient presenting with mild leukocytosis. A subsequent larger study highlighted pitfalls, including a higher risk of misclassifying atypical cells as plasma cells, monocytes, myelocytes, or blasts. Alferez et al.'s 2015 platform using color features and watershed transformation for CLL recognition achieved 80% accuracy and 98.6% specificity for five lymphoid cell types. Their 2016 expansion to 4,000 images achieved 98% overall accuracy for screening normal vs. abnormal lymphocytes and 91.23% for sub-classification into specific disease entities.
Digital pathology for the diagnosis and grading of lymphoid neoplasms in tissue represents a particularly important application area given the clinical consequences of misgrading. Follicular lymphoma (FL) is the lymphoma entity with the highest number of AI studies, driven by the inherent subjectivity of its WHO grading system. Current FL grading requires counting centroblasts (CB) per high-power field (HPF): Grade I (0-5 CB/HPF) and Grade II (6-15 CB/HPF) constitute low-risk disease, while Grade III (greater than 15 CB/HPF) is high-risk. The narrow ranges of these thresholds combined with subjective counting can dramatically affect assigned grade, prognosis, and treatment decisions. Published inter-observer variation can reach up to 41% lack of consensus among expert pathologists. Digital reading from WSI with preselected regions has been shown to improve inter-reader agreement, reducing disagreement to only 5.9% for centroblast enumeration.
Follicle detection using IHC: Accurate automated identification of neoplastic follicles is a prerequisite for appropriate high-power field selection for centroblast counting. Samsi et al. used iterative watershed and color and texture features combined with an unsupervised k-means clustering algorithm on 8-12 IHC images stained for B-cell markers (CD10, CD20), achieving 87% accuracy for follicle segmentation compared to manual segmentation. Oger et al. developed a region-based segmentation approach using CD20 to delineate follicles, then refined results by mapping follicle boundaries on high-resolution H&E images using a k-means classifier on 12 cases. Importantly, automated identification of all follicles in a tissue section enables centroblast counting across the entire available tissue rather than being limited to the conventional 10 HPF standard.
H&E-only follicle detection: Studies using only H&E images for follicle identification have generally shown lower accuracy. Belkacem-Boussaid used concavity index calculation and recursive watershed operations to reduce over-segmentation bias, achieving 78% accuracy for follicle detection. A separate publication by the same group developed an automated method for centroblast identification independent of follicle identification, combining geometric and texture feature extraction with a supervised quadratic discriminant analysis classifier on 436 images; the accuracy for distinguishing centroblasts from non-centroblasts was 82%.
Combined FLAGS grading system: The Ohio University team published the Follicular Lymphoma Grading System (FLAGS) in 2015, which automatically identifies more than 10 candidate fields suitable for grading using both H&E and CD20 stains in combination, followed by classification of fields into high or low grade based on centroblast counts via a k-nearest neighbor classifier on 20 slides, achieving 80% accuracy. Sertel et al.'s earlier 2009 approach used cytologic components and spatial distribution with color texture analysis, achieving 98.9% sensitivity and 98.7% specificity for detecting Grade III FL but requiring additional standardization for lower-grade entities. A follow-up study by the same group using unitone conversion and a Bayesian classifier on 100 images achieved 80.7% detection accuracy for centroblasts.
Diffuse large B-cell lymphoma (DLBCL) is the most common aggressive lymphoma and has been an important target for AI classification efforts. Cell-of-origin (COO) subclassification using immunohistochemistry (IHC) into germinal center (GC) and non-GC subtypes carries prognostic significance under R-CHOP therapy but historically showed poor concordance with gene expression profiling (GEP). Da Costa developed a machine learning J48 algorithm incorporating CD10, MUM1, FOXP1, and BCL-6 markers from an automated IHC classification platform; 91.6% of 475 cases were correctly classified as GC or non-GC, with high concordance with GEP results and prognostic significance.
Comparative ML algorithm performance for COO classification: A Mexican group evaluated multiple ML structures, including ANN and SVM, for classifying DLBCL using a combination of IHC antibodies from different published algorithms. Their optimized algorithm achieved 94% accuracy, 93% specificity, and 95% sensitivity, with high agreement with GEP. A Chinese study compared six IHC algorithms across 855 cases of de novo DLBCL-NOS treated with CHOP/R-CHOP, using SVM to evaluate concordance with GEP. The Choi and Visco-Young algorithms showed the highest concordance rates, suggesting their particular utility for separating GCB from non-GCB subtypes for clinical risk stratification.
Prognostic modeling for DLBCL: A European group extended traditional prognostic indices (IPI, revised IPI, NCCN-IPI) by incorporating more clinical variables into a stacking ensemble ML model trained on 5,173 cases from Nordic lymphoma registries (the Biccler et al. study). This model, made publicly available at lymphomapredictor.org, represents one of the largest ML prognostic tools for DLBCL developed from population-level data. Using IPI alone as a baseline, the ML model demonstrated improved outcome prediction, particularly in intermediate-risk patient groups where traditional scoring is least discriminating.
Treatment resistance prediction: Treatment resistance to R-CHOP has been investigated using 2D and 3D CT radiomic analysis with random forest (RF) and SVM methods. Lymph node sections from patients with known treatment resistance were contoured and segmented before analysis. The radiomic prediction models provided high accuracy for identifying treatment-resistant DLBCL, offering a potential pathway to avoiding toxicity from ineffective treatment regimens. A CNN-based lymphoma diagnostic model classifying cases into four categories (benign lymph node, DLBCL, Burkitt lymphoma, and SLL) from H&E images on 2,560 images achieved a diagnostic accuracy of 100%, representing the only study at the time to use DL for a multi-class lymphoma diagnostic classification task.
The advent of immunotherapy and targeted therapy has created an urgent clinical need to understand the co-expression patterns of multiple tumor cell markers and the interaction between tumor cells and their microenvironment. Conventional IHC is limited to detecting one or two markers per slide, making it difficult to characterize the spatial relationships between different immune and tumor cell populations. Multiplexed immunofluorescence (mIF) technologies that can simultaneously stain and image many markers on a single tissue section represent a transformative advance for this purpose, particularly in lymphoproliferative diseases where the complex cellular interactions within the lymph node microenvironment drive disease behavior.
Available multiplexing platforms: The review catalogues the major multiplexed staining technologies currently available. The MultiOmyx platform, belonging to the multiplex staining bleaching category, allows analysis of up to 60 biomarkers on a single tissue section. The Multiplexed Ion Beam Imaging (MIBI) technology, based on mass spectrometry imaging principles, can analyze up to 100 biomarkers simultaneously, though with a longer staining time. Additional platforms include Multiplex signal amplification techniques (such as Opal/TSA-based systems) and various MALDI-TOF mass spectrometry imaging approaches. Automated scanning products commercially available include AxioVision MosaiX and others capable of acquiring the multi-channel images necessary for downstream AI analysis.
Clinical applications in hematopathology: For lymphoproliferative diseases, mIF holds particular promise for characterizing the tumor microenvironment of DLBCL and FL, identifying spatial patterns of T-cell infiltration, immune checkpoint expression, and co-stimulatory molecule patterns that may predict response to immunotherapy. Pinpointing small micro-metastases in lymph nodes, which can be challenging on H&E alone, would be facilitated by simultaneous immunostaining of multiple lymph node components to highlight any abnormality. The integration of AI-based image analysis algorithms for multi-channel mIF data is an active and rapidly evolving area. The quality and consistency of staining and scanning are emphasized as critical prerequisites, particularly when tissue is limited to small core biopsies.
The authors position mIF as one of the most important future directions for AI-assisted hematopathology, given its ability to generate rich spatial data about biomarker co-expression that is ideally suited to pattern recognition algorithms. The ability to profile 60-100 biomarkers in a single tissue section, combined with AI analysis, has the potential to discover entirely new prognostic and predictive signatures in lymphoma.
Scope of current limitations: The authors identify several key barriers constraining the clinical deployment of AI in hematopathology. Most studies are based on small, single-center datasets with a limited number of images or cases. This is particularly problematic for hematologic malignancies, where individual subtypes are rare and assembling large training cohorts requires multi-institutional collaboration. Many models are trained on straightforward cases without including morphologically complex or borderline entities, meaning published accuracy figures likely overestimate real-world performance on the full spectrum of diagnostic cases encountered in clinical practice.
Technical standardization gaps: Variability in image acquisition, including differences in scanner hardware, scanning parameters, staining protocols, and tissue preparation, introduces domain shift that degrades model performance when deployed at institutions other than the one where training data was collected. The lack of standardization in image acquisition was specifically highlighted in the largest deep learning survey to date, encompassing more than 84 studies across nine cohorts, as a critical prerequisite for breakthrough of DP models into clinical practice. For hematology specifically, WSI limitations such as the unavailability of three-dimensional images present additional technical constraints not encountered in solid tumor pathology.
The pathologist-AI relationship: The review is explicit that the outcomes of any digitalization system will ultimately require review and supervision by a pathologist, who must approve or reject machine-derived results while considering histologic findings, clinical presentation, and contingent factors. The primary aim of digital pathology should be to facilitate and standardize the diagnostic process to complement human expertise, not to replace it. This framing is important both conceptually and practically, since regulatory approval pathways for AI diagnostic tools require demonstration of performance within a supervised workflow rather than as autonomous decision-makers.
Priority future directions: The authors point to several high-priority areas: scaling datasets through federated multi-institutional collaboration; moving beyond morphology-only models toward multimodal AI that integrates digital pathology, flow cytometry, cytogenetics, and molecular data in a unified framework; developing mIF-based AI tools that leverage spatial biomarker co-expression for lymphoma subtyping and microenvironment characterization; prospective validation studies embedded in real clinical workflows; and leveraging unsupervised learning methods to reduce the annotation burden that has historically limited labeled dataset construction in hematopathology.