B-cell acute lymphoblastic leukemia (B-ALL) is the most common cancer in children, accounting for 80% of all childhood leukemias. It arises when immature B-lymphocytes (a type of white blood cell) become cancerous and multiply uncontrollably in the bone marrow. Modern treatment with multi-agent chemotherapy has been remarkably successful, achieving 5-year survival rates exceeding 90%.
Despite this success, approximately 20% of children with B-ALL will experience a relapse - a return of the leukemia after initial treatment. This occurs because a small population of cancer cells survives the initial chemotherapy, often having developed mechanisms of drug resistance. Relapsed B-ALL is far more difficult to treat than newly diagnosed disease.
The prognosis for relapsed B-ALL is severe. Even with intensified chemotherapy or stem cell transplantation - the most aggressive treatment options available - overall survival rates for relapsed patients are only 35-40%. Relapsed B-ALL has become the leading cause of cancer-related deaths among children, even as outcomes for newly diagnosed leukemia have dramatically improved.
A major obstacle to improving outcomes is that the biological mechanisms driving relapse are still poorly understood. What changes in the cancer cells between diagnosis and relapse? What mutations or gene expression changes make them resistant? Answering these questions requires analyzing matched samples from the same patients at both time points.
Several research groups have performed gene expression microarray studies - technology that measures the activity of thousands of genes simultaneously - comparing cancer cells at diagnosis versus at relapse in the same patients. This paired design is ideal for identifying which gene changes specifically occur during the process of relapse.
However, these individual studies have produced discordant and unreliable results. When the same statistical methods were applied to reanalyze three existing studies, virtually no genes were consistently altered across all three. The main reason is small sample sizes: each study included fewer than 50 patients, giving insufficient statistical power to reliably detect the moderate but real gene expression changes associated with relapse.
Different studies also used different microarray platforms, different normalization methods, and different statistical approaches - making direct comparisons across studies unreliable. The result is that the published literature contains lists of candidate genes that differ substantially from study to study, making it difficult to identify which changes are genuinely important and which are statistical noise.
Meta-analysis - the systematic combination of data from multiple studies - offers a solution. By pooling data from all available studies and applying a unified statistical framework, meta-analysis dramatically increases statistical power and can identify reproducible signals that no individual study had the power to detect reliably.
The researchers searched the Gene Expression Omnibus (GEO) and ArrayExpress databases - public repositories of microarray datasets - using the terms 'acute lymphoblastic leukemia' and 'microarray.' From over 550 relevant datasets, only three met the strict inclusion criteria: having paired diagnosis-relapse samples from the same pediatric B-ALL patients with available raw data files.
These three datasets (GSE18497, GSE28460, and GSE3910) collectively contained 108 matched diagnosis-relapse sample pairs - far more than any individual study. GSE3910 used an older Affymetrix array platform, while GSE18497 and GSE28460 used the more comprehensive Affymetrix Human Genome U133 Plus 2.0 array. Technical differences between platforms were addressed using batch correction with the COMBAT algorithm.
The statistical method chosen was RankProd - a non-parametric (distribution-free) meta-analysis approach. Rather than assuming gene expression values follow a specific statistical distribution, RankProd ranks genes by how consistently they appear highly up- or down-regulated across multiple datasets. This approach is robust to platform differences and small sample sizes, and performs better than competing methods at identifying reproducibly differentially expressed genes.
The significance threshold used was a percentage of false positives (pfp) below 1% - estimated using 1,000 random permutations of the data. This is equivalent to a false discovery rate of less than 1%, meaning fewer than 1 in 100 of the reported gene changes is expected to be a statistical artifact. Genes were additionally required to show a fold change greater than 1 between relapse and diagnosis.
The meta-analysis identified dramatically more differentially expressed genes than any individual study had found. 1,527 genes were significantly upregulated at relapse compared to initial diagnosis, and 1,214 genes were significantly downregulated, all meeting the strict false discovery rate threshold. By comparison, individual analysis of the three datasets separately found only 1-23 significantly upregulated genes each, with almost no overlap between studies.
This dramatic improvement in discovery power illustrates the fundamental advantage of meta-analysis: combining 108 patient pairs versus 27-49 in individual studies yields the statistical power needed to detect real but moderate-sized effects. The top-ranked genes showed extremely high statistical confidence (pfp = 0, meaning none of the 1,000 permutation tests produced a more extreme result).
S100A8 emerged as the most significantly upregulated gene at relapse, with a fold change of 1.87 - meaning it was expressed nearly twice as highly at relapse as at diagnosis. This was not a new finding from the meta-analysis alone: S100A8 was also identified as upregulated in 2 of the 3 individual studies, making it one of only two genes to show consistent upregulation across individual analyses. The meta-analysis confirmed and strengthened this signal.
Hierarchical clustering of the top 100 differentially expressed genes revealed that diagnosis and relapse samples were not perfectly separated by their gene expression profiles - they were mixed rather than forming distinct clusters. This finding suggests that the molecular differences between diagnosis and relapse are subtle and gradual, consistent with the idea that relapse-driving cells are already present at a low frequency at diagnosis rather than arising de novo after treatment.
Gene ontology analysis of the 1,527 upregulated genes revealed that the most significantly enriched biological category was cell cycle processes, with an enrichment score of 15.3. Among the top 100 upregulated genes, 14 were direct cell cycle regulators that also interact with each other in a protein-protein interaction network - suggesting a coordinated rewiring of cell cycle control at relapse.
The key upregulated cell cycle genes included CDK1 (cyclin-dependent kinase 1, which governs passage through cell division), AURKA (Aurora kinase A, which controls the formation of the mitotic spindle and centrosome maturation), BIRC5/survivin (which suppresses cell death during division), BUB1B and TTK (spindle assembly checkpoint proteins), and multiple kinesin motor proteins that drive chromosome separation.
The 1,214 downregulated genes at relapse were most enriched in transcription regulation pathways (enrichment score = 12.6). This suggests that relapsed leukemia cells have altered gene regulatory networks - potentially losing the normal controls that limit cell proliferation or enforce cell identity. Reduced expression of transcription factors could contribute to the loss of differentiation capacity seen in aggressive leukemia.
The consistent overactivation of cell cycle machinery at relapse across all three datasets points to a fundamental biological mechanism: relapsed ALL cells may survive chemotherapy precisely because they proliferate faster and more aggressively, outcompeting attempts to kill them with drugs designed to target rapidly dividing cells. Paradoxically, this means targeting the cell cycle machinery itself may be the most direct therapeutic approach.
S100A8 is a small calcium-binding protein that belongs to the S100 gene family. It is normally expressed in neutrophils (a type of white blood cell) and plays roles in inflammation. In cancer, it acts as an oncogenic factor promoting cell proliferation, invasion, and resistance to chemotherapy-induced cell death.
In leukemia specifically, S100A8 has been found overexpressed in childhood AML patients with worse prognosis. In B-ALL, it appears especially elevated in aggressive subtypes - notably infant B-ALL and the MLL-rearranged subtype, both of which carry poor prognoses. High S100A8 expression has been linked to prednisolone (steroid) resistance in infant ALL, directly connecting it to treatment failure.
S100A8 may promote drug resistance partly by stimulating autophagy - a cellular self-digestion process - which allows leukemia cells to survive stress caused by chemotherapy. It also activates the MAPK and NF-kB signaling pathways that support cell survival and proliferation. Silencing S100A8 with RNA interference (siRNA) has been shown to increase drug sensitivity and trigger cell death in leukemia cells.
Several existing drugs could potentially target S100A8 indirectly. Since S100A8 acts as an upstream activator of EGFR signaling, drugs targeting EGFR (such as gefitinib) may be relevant. The MAPK inhibitor SB203580 and the NF-kB inhibitor Bay 11-7082 have both shown ability to block S100A8-mediated cancer cell invasion in laboratory studies. The authors propose S100A8 as the most promising single gene target for further validation studies in relapsed B-ALL.
CDK1 (cyclin-dependent kinase 1) orchestrates the final steps of cell division and is overexpressed at relapse. CDK inhibitors as a class have already shown promising activity in leukemia: a clinical trial combining the CDK inhibitor flavopiridol with standard chemotherapy achieved a 75% complete remission rate in AML patients, compared to 40-50% with chemotherapy alone. The CDK inhibitor dinaciclib has shown significant activity in relapsed chronic lymphocytic leukemia (CLL). Palbociclib was FDA-approved for breast cancer in 2015. Whether CDK inhibitors can specifically benefit children with relapsed B-ALL is still under investigation.
AURKA (Aurora kinase A) was also prominently upregulated at relapse. Aurora kinases regulate cell division by controlling centrosome function and spindle assembly - the machinery that pulls chromosomes apart during cell division. Overexpression of AURKA correlates with higher tumor grade and worse prognosis across many cancer types. Over 30 Aurora kinase inhibitors have been tested clinically. For relapsed AML, the AURKA inhibitor MLN8237 (alisertib) showed 13% complete response and 49% stable disease in an early phase I/II trial.
Survivin (BIRC5) is both a cell cycle regulator and an anti-apoptotic protein. It forms part of the chromosomal passenger complex essential for cell division and simultaneously blocks programmed cell death. Survivin overexpression has already been identified as a strong prognostic risk factor for relapse in childhood B-ALL by prior individual studies, and its upregulation was confirmed and strengthened by this meta-analysis. Multiple clinical trials targeting survivin using antisense oligonucleotides, small molecule inhibitors, and immunotherapy approaches are ongoing.
The interconnected network of cell cycle genes identified by this analysis is particularly important: CDK1, AURKA, and survivin are not independent targets but form part of an integrated signaling complex. Combinations of inhibitors targeting multiple nodes in this network may be more effective than single-target approaches and warrant pre-clinical investigation specifically in relapsed B-ALL models.
The meta-analysis demonstrates that combining small, existing datasets can yield robust, reproducible findings that individual studies cannot achieve on their own. The identification of over 2,700 consistently dysregulated genes provides a substantially more comprehensive picture of relapse biology than was previously available, and creates a rich resource for hypothesis generation.
The authors acknowledge limitations: the total sample size of 108, while larger than any individual study, remains modest. The study also relies on data from archived datasets generated with different microarray technologies, which introduces technical variability even after batch correction. Prospective validation in independent patient cohorts with larger numbers will be needed before any biomarkers can be used clinically.
For S100A8 specifically, the next step should be functional studies in leukemia cell lines and animal models to confirm that inhibiting S100A8 can re-sensitize relapsed B-ALL cells to standard chemotherapy. If confirmed, clinical translation would involve testing whether existing drugs that indirectly inhibit S100A8 signaling could be combined with salvage chemotherapy in relapsed patients.
The broader message of this work is that the bottleneck in understanding relapse is not lack of data collection - three independent studies had already collected matched samples - but rather the statistical power to analyze them. The meta-analysis approach should be applied more systematically to other pediatric cancers where individual studies have failed to produce consistent findings.