Low-Cost Transcriptional Diagnostic to Accurately Categorize Lymphomas in LMICs

Blood Advances 2021 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
The Diagnostic Crisis Facing Lymphoma Patients in Low- and Middle-Income Countries

Accurate lymphoma diagnosis in high-income countries (HICs) requires a pathologist to integrate hematoxylin and eosin (H&E) histology, immunohistochemistry (IHC), flow cytometry, cytogenetics, and molecular profiling. The World Health Organization (WHO) lists histopathology with IHC among the essential in vitro diagnostics for facilities with clinical laboratories. Yet in low- and middle-income countries (LMICs), this standard of care is largely inaccessible - both because of the cost and because of a severe shortage of trained pathologists.

Scale of the problem: At the Instituto de Cancerologia y Hospital Dr. Bernardo Del Valle (INCAN) in Guatemala City, the country's only public cancer hospital serving both urban and rural indigenous Mayan communities, approximately 100 patients per year present with findings suspicious for lymphoma. Limited IHC is available at an out-of-pocket cost of approximately $450 per patient - a sum beyond the reach of most patients. As a result, most biopsy specimens are evaluated only by H&E staining, yielding ambiguous diagnoses such as "suspect large cell lymphoma" or "lymphoma-not otherwise specified (NOS)," which cannot guide targeted treatment selection.

Why this matters clinically: Many lymphoma subtypes are curable or highly responsive to treatment regimens that are accessible within LMICs, including CHOP-based chemotherapy for DLBCL and follicular lymphoma, ABVD for Hodgkin lymphoma, and SMILE for natural killer/T-cell lymphoma (NKTCL). The barrier is not the absence of effective treatment - it is the absence of accurate diagnosis to direct that treatment. This is a situation analogous to the HIV crisis in LMICs during the late 1990s, when the deployment of highly active antiretroviral therapy required the parallel development of inexpensive, operator-independent diagnostics.

This study from Valvert et al., published in Blood Advances in 2021, directly addresses this gap by developing and validating a low-cost, machine learning-based transcriptional assay capable of classifying lymphoma subtypes from formalin-fixed, paraffin-embedded (FFPE) biopsy specimens without requiring prior pathologist review - at a reagent and consumable cost of approximately $10 per sample.

TL;DR: Lymphoma diagnosis in LMICs is critically limited by the ~$450 cost of IHC and pathologist shortages. INCAN Guatemala sees ~100 suspected lymphoma cases per year but most receive only H&E staining with ambiguous results. This paper develops a ~$10/sample gene expression assay plus machine learning classification to replace expert pathology for lymphoma subtyping.
Pages 2-3
Building a 13-Year Biopsy Cohort and Establishing Ground-Truth Diagnoses

The study reviewed all biopsy specimens collected at INCAN between 2006 and 2018 for clinical suspicion of lymphoma, identifying 3,015 tissue blocks from 1,836 individual patients. After excluding core biopsies and blocks smaller than 1 cm3 to preserve tissue for future clinical needs, 670 banked FFPE samples were included: 650 obtained for suspected lymphoma and 20 from normal tonsillar tissue serving as benign controls. Clinical data were extracted by manual review of paper charts - a labor-intensive process reflecting the resource constraints typical of LMIC settings.

WHO-based pathology as the gold standard: Half of each FFPE block was shipped to Stanford University, where H&E slides were generated from whole sections and reviewed by two expert hematopathologists (O.S. and Y.N.) blinded to original INCAN diagnoses. Tissue microarrays (TMAs) were constructed from two representative cores per case and subjected to IHC using more than 30,000 individual stain assessments. Epstein-Barr virus (EBV) was evaluated by IHC for EBV-LMP1 and in situ hybridization. Fluorescence in situ hybridization (FISH) was performed for MYC, BCL2, and BCL6 rearrangements. All cases were classified according to the 2016 WHO classification.

Diagnostic binning: Rather than attempting to reproduce the full WHO classification granularity, diagnoses were collapsed into 9 therapeutically driven categories: (1) aggressive B-cell lymphoma including Burkitt and B-lymphoblastic; (2) diffuse large B-cell lymphoma (DLBCL); (3) Hodgkin lymphoma (HL); (4) marginal zone lymphoma (MZL); (5) mantle cell lymphoma (MCL); (6) follicular lymphoma (FL); (7) natural killer/T-cell lymphoma (NKTCL); (8) T-cell lymphoma (TCL); and (9) nonmalignant. This treatment-directed framing is a key design decision - the assay was optimized for guiding clinical management rather than achieving maximum taxonomic precision.

TMA validation: To verify that TMA cores accurately represented the full biopsy, 80 cases were randomly selected with stratified sampling proportional to diagnostic bin frequency. Blinded rereview of whole sections after a one-year washout was concordant with TMA diagnoses in 77 of 78 cases (98.7%), confirming the validity of the TMA-based diagnostic reference standard.

TL;DR: 670 FFPE biopsy specimens from a 13-year INCAN Guatemala cohort. Expert Stanford hematopathologists established ground-truth diagnoses via WHO 2016 classification using H&E, IHC (30,000+ stains), EBV testing, and FISH. Diagnoses were binned into 9 treatment-directed categories. TMA-to-whole-section concordance was 98.7% (77/78 cases).
Pages 3-4
The Chemical Ligation Probe-Based Assay: 37 Genes at $10 Per Sample

The core innovation of this study is the chemical ligation probe-based assay (CLPA), which quantifies the expression of 37 genes plus 2 control genes directly from FFPE tissue scrolls using standard capillary electrophoresis equipment. The 37 genes were selected from prior publications based on three criteria: lineage- or subtype-specific expression, prognostic value, or therapeutic relevance. The gene panel includes well-established lymphoma markers such as CD5, CD20 (MS4A1), BCL2, BCL6, CCND1, PAX5, IRF4, MKI67, MYC, ALK, EBER1, and NCAM1, along with others associated with specific TCL subtypes (SOX8, GATA3, TBX21, ICOS).

Assay protocol: Ten-micrometer scrolls from paraffin-embedded tissue blocks were cut and processed without prereview for minimum tumor cell content. Scrolls were combined with ligation reaction buffers (DxBuffer1, Lymphoma RUO Mix A and B, DirectMix C) in a 96-well PCR plate and ligated at sequential temperatures (55°C, 80°C, 55°C). Following magnetic bead cleanup, PCR amplification was performed (30 cycles, hot start at 95°C), and products were analyzed by capillary electrophoresis on an Applied Biosystems 3500 or SeqStudio Genetic Analyzer. Two normalizer genes (ISY1 and WDR55) expressed in all fluorescent channels were included for quantitative normalization.

Cost structure: Based on actual supply purchasing experience in Guatemala, manufacturing of assay-specific reagents costs approximately $5 per sample. Running 95 samples per week with appropriate controls brings the total reagent plus consumable cost to $6.76 per sample. Even at lower throughput (16 tests per week), the cost does not exceed $15 per sample - compared to $450 for limited IHC at INCAN. This 30- to 60-fold cost reduction is the central health-economic argument for the assay.

Quality control: Ligation control (LIGC) and PCR control (PCRC) genes were included in each reaction to flag assay failures. Of the 670 samples, 60 (8.9%) failed QC - predominantly older specimens. Among biopsy specimens from 2015 or later (less than 3 years old at time of testing), only 4 of 194 (2%) failed QC, demonstrating that failure rates decline substantially with fresher tissue.

TL;DR: The CLPA quantifies 37 lymphoma-relevant genes from FFPE scrolls by capillary electrophoresis at a cost of $6.76-$15 per sample depending on throughput, versus $450 for limited IHC. No pathologist prereview is needed. QC failure rate is 8.9% overall but only 2% for specimens less than 3 years old. Normalizer genes ISY1 and WDR55 enable quantitative expression measurement.
Pages 4-5
Super Learner Ensemble: 13 Base Models Feeding an XGBoost Classifier

After quality control filtering, the 560 samples obtained before treatment were split into training (70%, n=397) and validation (30%, n=163) cohorts, with stratified sampling to maintain proportional representation of all 9 diagnostic bins. The 39 post-relapse biopsy specimens formed a separate test set. Gene expression values were normalized by subtracting the mean of log2 normalizer signals from the log2 signal of each response gene. Features were centered and scaled using preprocessing parameters derived exclusively from the training set.

Base learners: Thirteen candidate classification models were evaluated and tuned on the training set using five repeats of 10-fold cross-validation with logarithmic loss as the performance metric. The candidate models spanned diverse algorithmic families: linear discriminant analysis (LD1, LD2), mixture discriminant analysis (MDA), penalized multinomial regression (MULTINOM), penalized discriminant analysis (PDA), partial least squares (PLS), naive Bayes (NB), neural networks (NN), k-nearest neighbors (KNN), nearest shrunken centroids (PAM), high-dimensional discriminant analysis (HDDA), support vector machines with polynomial kernel (SVMPOLY) and radial kernel (SVMRAD), and random forest (RF). Grid search and random search optimization were used to tune hyperparameters.

Super learner architecture: Class probabilities from all 13 base learner models were concatenated and used as input features for an extreme gradient boosting (XGBoost) super learner model. This model stacking approach leverages the diversity of the individual base learners - each of which captures different aspects of the gene expression data - and lets the XGBoost meta-learner learn which combination of predictions is most reliable for each diagnostic bin. The final output is a probability distribution across all 9 bins, with the bin receiving the highest probability assigned as the predicted class.

Indeterminate calls: Cases where the maximum class probability was below 60% were classified as indeterminate rather than assigned to any bin. This threshold was chosen conservatively to prioritize specificity in high-confidence calls. In the validation cohort, 28 of 163 cases (17%) were classified as indeterminate. Among the 135 with a probability of at least 60%, 83% had probabilities above 90%, indicating that the model is highly confident when it does make a call.

TL;DR: 13 base learners (LDA, MDA, MULTINOM, PDA, PLS, NB, NN, KNN, PAM, HDDA, SVMPOLY, SVMRAD, RF) were tuned by 5-repeat 10-fold cross-validation. Class probabilities fed an XGBoost super learner. Calls with less than 60% confidence were classified as indeterminate (17% of validation cases). Of high-confidence calls, 83% had greater than 90% probability.
Pages 5-6
86% Overall Accuracy, 94% After Excluding Indeterminate Calls

In the 163-sample validation cohort, the CLPA achieved an overall accuracy of 86% (95% CI: 80%-91%) against the IHC-based gold standard. Performance varied by lymphoma subtype: DLBCL, HL, MCL, and NKTCL each achieved at least 90% accuracy, while FL (77%), MZL (60%), and aggressive B-cell lymphoma (50%) showed lower performance, likely reflecting the smaller training set size for those subtypes and their greater gene expression overlap with other bins.

High-confidence call accuracy: After excluding the 28 indeterminate cases (17% of the cohort), accuracy among the 135 high-probability calls (probability at least 60%) increased to 94% (95% CI: 89%-97%). Only 8 of 135 confident calls were misclassified. The specific misclassifications were: 1 case called DLBCL by CLPA but B-lymphoblastic lymphoma by IHC; 2 cases called FL but actually DLBCL; 2 cases called DLBCL but actually FL; 2 cases called HL but actually anaplastic large cell lymphoma or PTCL-NOS; and 1 case called NKTCL but actually PTCL-NOS. Notably, in one of the "misclassified" HL vs. PTCL-NOS cases, subsequent rereview of the whole section revealed rare CD30+, CD15+ (subset), PAX5+ (variable), MUM1+, EBV-negative cells consistent with classical HL, indicating that the CLPA had actually made the correct call and the IHC diagnosis was the error.

Nonmalignant cases: Approximately 7.5% of biopsies performed for suspected lymphoma at INCAN were classified as nonmalignant by expert IHC, and many of these patients had received cytotoxic chemotherapy based on erroneous INCAN diagnoses. The CLPA classified all nonmalignant cases as either nonmalignant or indeterminate - with zero false positives misclassifying a benign case as malignant lymphoma in the high-confidence call set. Conversely, only 1 of 234 malignant lymphoma cases across the three validation cohorts was called nonmalignant with high confidence by the CLPA.

Relapsed/refractory cohort: Forty-two samples from patients with relapsed or refractory disease were assessed, with 39 evaluable. Overall accuracy was 79% (95% CI: 64%-91%), increasing to 88% (95% CI: 73%-97%) after excluding 5 indeterminate cases. The slightly lower performance relative to diagnostic specimens likely reflects increased biologic heterogeneity and potential subtype evolution at relapse.

TL;DR: Validation cohort accuracy: 86% overall (163 cases), 94% among high-confidence calls (135 cases, 95% CI 89%-97%). DLBCL, HL, MCL, and NKTCL each exceeded 90% accuracy. All nonmalignant cases were correctly called nonmalignant or indeterminate - zero false malignant calls. Relapsed cohort accuracy: 79% overall, 88% after excluding indeterminates.
Pages 6-7
97% Concordance Between Guatemala City and California Laboratories

A critical practical question for deploying this assay in LMICs is whether it can be run reliably by laboratory staff at the site of care rather than at a sophisticated reference laboratory. To address this, the researchers conducted a prespecified multisite comparison using 59 cases randomly selected from the validation cohort in proportion to lymphoma subtype incidence. Consecutive 6-micrometer sections were cut from each block and either processed at INCAN in Guatemala City or shipped to DxTerity Diagnostics in Rancho Dominguez, California. Technicians at each site ran the assay blinded to the other site's results.

Results: Of the 59 cases, 58 passed QC at both sites (the single failure failed at both sites). Among the 58 evaluable pairs, 37 reached the 60% probability diagnostic threshold at both sites simultaneously. The diagnosis was concordant in 36 of 37 of these high-confidence pairs (97%). The single discordant call was classified as FL by INCAN and DLBCL by DxTerity. In 9 additional cases, only one site reached the 60% threshold; 8 of those 9 were concordant with IHC-based diagnosis. The remaining 12 pairs were indeterminate at both sites.

This 97% concordance rate for high-confidence calls across two laboratories - one in a well-resourced commercial setting in the United States and one in a public cancer hospital in Guatemala City - provides strong evidence that the assay is robust to site-to-site variability in laboratory technique. The fact that INCAN laboratory staff could run the assay independently after training is particularly important for scalability: the assay does not require shipping specimens abroad for analysis, and INCAN has demonstrated it can turn around results in less than 24 hours when urgent.

Broader implications: Based on experience with the Max Foundation's low-cost BCR-ABL GeneXpert assay for chronic myelogenous leukemia deployed across more than 70 countries, the authors note that most large LMIC cancer centers have access to PCR machines and capillary electrophoresis instruments - making the CLPA deployable without major capital investment at centers that already manage hematologic malignancies.

TL;DR: In a blinded multisite comparison, INCAN Guatemala and DxTerity California produced concordant diagnoses in 36 of 37 high-confidence paired calls (97%). QC concordance was 58/59. The assay is deployable by local LMIC laboratory staff without specimen shipment. Turnaround under 24 hours is achievable at INCAN.
Pages 7-8
Performance Gaps, Diagnostic Boundary Cases, and Benchmarking Challenges

Subtype-specific limitations: The assay's accuracy is not uniform across all 9 diagnostic bins. Performance was lowest for aggressive B-cell lymphoma (50% for the 2 validation cases, though this small n limits interpretation), MZL (60%), and FL (77%). MZL and FL share considerable gene expression overlap and are difficult to distinguish even by expert IHC - discordance rates between expert hematopathologists for these entities exceed those for more morphologically distinct subtypes. The training set was also small for aggressive B-cell lymphoma (7 training cases) and MZL (13 training cases), limiting the model's ability to learn subtype-specific features for these less common entities at INCAN.

The benchmarking problem: The accuracy figures cited above use IHC-based expert diagnosis as the reference standard, but clinical pathology is itself an imperfect benchmark. Discordance rates between expert hematopathologists in the diagnosis of lymphoma are approximately 10% in HICs and exceed 30% in LMICs. The authors note that at least one apparent CLPA error was subsequently shown to be a correct call after rereview of the whole section. Gene expression studies have also consistently shown that a fraction of lymphomas classified as one entity by IHC cluster more closely with a different entity in molecular space (for example, Burkitt-like DLBCLs). True CLPA accuracy may therefore be higher than the reported figures suggest.

Burkitt lymphoma and B-lymphoblastic lymphoma binning: These two biologically and therapeutically distinct entities were combined into a single "aggressive B-cell" bin, a simplification justified by their rarity at INCAN (only 10 combined cases over 12 years) but one that leaves a clinically significant ambiguity. The authors acknowledge that when this bin is called, H&E and terminal deoxynucleotide transferase (TdT) IHC staining would still be required to distinguish between the two diagnoses.

Indeterminate rate and out-of-bin diagnoses: Approximately 15-20% of all cases, including those with diagnoses outside the 9 bins (CLL/SLL, plasma cell neoplasms, T-lymphoblastic lymphoma, plasmablastic lymphoma, carcinoma), would require additional workup including IHC. Among the 32 out-of-bin cases tested, 78% were correctly flagged as indeterminate. This means roughly 22% of such cases received a confident but potentially misleading bin assignment - a limitation with real clinical consequences if clinicians interpret all CLPA calls as definitive.

TL;DR: Accuracy is lower for MZL (60%), FL (77%), and rare subtypes with small training sets. Expert IHC itself has 10-30% discordance, making the 86-94% CLPA figures potentially understated. Burkitt and B-lymphoblastic lymphoma remain combined in one bin. 17% of validation cases were indeterminate; 22% of out-of-bin diagnoses received a confident (potentially incorrect) call.
Pages 8-9
Prospective Expansion, Assay Redesign, and Open-Access Deployment

Prospective multicenter study: The authors initiated a prospective study of CLPA-based diagnosis across centers in Guatemala, El Salvador, and Belize, with all testing performed centrally at INCAN. This represents the next critical validation step - moving from retrospective banked specimen analysis to real-time prospective classification of newly diagnosed patients, with the opportunity to track concordance with clinical outcomes and treatment responses.

Assay expansion: The current chemistry using standard capillary electrophoresis equipment can be extended from 37 to 55 genes without requiring new instrumentation. The authors are redesigning the assay to separately distinguish Burkitt lymphoma, B-cell lymphoblastic lymphoma, and plasma cell neoplasms - three entities currently either combined into one bin or flagged as indeterminate. Additional cohorts from other LMICs with different genetic backgrounds, coexisting infectious pathology (tuberculosis, HIV), and different subtype frequencies (for example, higher rates of NKTCL in Asian populations) are planned to improve generalizability and subtype-specific performance.

Needle biopsy compatibility: Current training and validation was performed on single 10-micrometer scrolls from excisional biopsy specimens ranging from 5x10 mm to 20x20 mm in surface area - corresponding to tissue volumes of 0.5 to 4.0 mm3. This tissue quantity is comparable to half or less of a core needle biopsy, suggesting the assay may be compatible with needle biopsies, which are frequently used in resource-limited settings. Formal testing on core needle biopsy specimens is ongoing.

Open-access R Shiny application: The research team is building an open-access R Shiny web application that accepts CLPA gene expression input data and returns a diagnosis and probability score from the machine learning model. This would allow any laboratory anywhere in the world with CLPA data to access the classification model without requiring local bioinformatics expertise. The authors also explicitly note that analogous gene panels and calling algorithms can be developed for other cancer types, positioning the CLPA framework as a generalizable platform for low-cost molecular diagnostics across oncology in resource-limited settings.

TL;DR: Active prospective study spans Guatemala, El Salvador, and Belize. Assay expansion to 55 genes will separately distinguish Burkitt, B-lymphoblastic, and plasma cell diagnoses. Needle biopsy compatibility is under investigation. An open-access R Shiny app is planned to enable any LMIC lab with CLPA data to receive ML-based diagnoses globally without local bioinformatics infrastructure.