Acute myeloid leukemia (AML) is the most common form of acute leukemia in adults, and its long-term survival rate remains poor across the general patient population. The disease arises when the bone marrow produces abnormal blood cells that crowd out healthy ones, leading to life-threatening complications.
A pivotal milestone in AML treatment is achieving complete remission (CR) - a state in which no leukemia cells can be detected in blood or bone marrow after chemotherapy. Patients who reach complete remission have significantly better survival outcomes than those with refractory (treatment-resistant) disease.
For patients who achieve complete remission and face intermediate or high-risk disease, allogeneic stem cell transplantation - replacing diseased bone marrow with healthy donor marrow - offers a potentially curative option. However, when disease fails to respond to treatment, outcomes are extremely poor even after transplantation.
Identifying which patients are likely to respond to intensive induction chemotherapy - and which are not - is critically important for tailoring treatment. Traditional prediction tools based on genetic and clinical markers exist but have moderate accuracy, motivating researchers to explore newer approaches like machine learning (ML).
Machine learning (ML) is a branch of artificial intelligence in which computers learn patterns from large datasets to make predictions - without requiring researchers to manually specify hypotheses in advance. This data-driven approach can potentially uncover relationships that traditional statistical models miss.
In this study, researchers applied nine different ML algorithms to data from 1,383 AML patients treated at multiple medical centers across Germany. The goal was to predict two critical outcomes: whether a patient would achieve complete remission, and whether they would survive two or more years after diagnosis.
The ML models incorporated 212 variables per patient, including clinical measurements (age, blood counts), laboratory values (hemoglobin, lactate dehydrogenase), chromosome abnormalities (cytogenetics), and gene mutations (molecular genetics). This multimodal approach reflects the complexity of AML biology.
Unlike conventional models that start with a pre-specified hypothesis, the ML algorithms automatically selected the most informative features from the full dataset. This allowed them to independently identify which patient characteristics matter most for predicting treatment outcome.
The study drew from four multicenter German clinical trials and a national AML registry spanning 59 specialized centers. All 1,383 patients were adults newly diagnosed with AML who received intensive induction chemotherapy. Patients with acute promyelocytic leukemia (APL) - a distinct subtype with very different biology and treatment - were excluded.
The machine learning pipeline involved several automated steps: first removing rare or sparse variables, then normalizing data, filling in missing values through statistical imputation, and running five separate feature-selection algorithms to determine which variables carried real predictive weight. Only variables supported by a defined threshold of importance were included in the models.
Nine ML classifier types were trained and tested, including random forest, gradient boosting, logistic regression, support vector machines (SVM), and artificial neural networks. Each algorithm learns patterns differently, which is why using multiple classifiers is valuable - no single method is best in all situations.
To prevent overfitting - where a model memorizes training data instead of learning general patterns - all test data were strictly withheld during the training phase. The models were then externally validated on a separate cohort of 664 AML patients from the German AML Cooperative Group (AMLCG) trials, who had not been seen by any algorithm during training.
Among the 1,383 patients, 72.9% achieved complete remission after intensive induction therapy. The ML models predicted whether a patient would reach CR with an accuracy - measured by area under the ROC curve (AUROC), where 1.0 is perfect and 0.5 is random - ranging from 0.77 to 0.86 in internal testing. Random forest performed best at 0.86.
The algorithms automatically selected 27 features as most informative for CR prediction. Patient age at diagnosis was the single most important predictor - each additional year of age was associated with a roughly 5.7% reduction in odds of achieving remission. Younger patients respond better, likely due to better tolerance of intensive chemotherapy and more favorable disease biology.
On the favorable side, mutations in NPM1, double-mutated CEBPA, chromosomal rearrangements t(8;21) and inv(16), and a normal karyotype were associated with higher odds of achieving complete remission. These are established markers of good-prognosis AML.
Conversely, mutations in TP53, ASXL1, RUNX1, U2AF1, SF3B1, and IKZF1, as well as chromosomal deletions del(5q) and del(17p) and the presence of extramedullary disease (leukemia cells spreading beyond the bone marrow), were strongly associated with failure to achieve remission. These are high-risk disease features.
Notably, the effect of FLT3-ITD mutations - a common and often feared AML mutation - depended on the proportion of mutated cells and concurrent NPM1 status. Low-level FLT3-ITD with NPM1 mutation was associated with better CR rates, while high-level FLT3-ITD reduced the odds of remission regardless of NPM1.
The same ML pipeline was applied to predict whether patients would survive two or more years after diagnosis. At the study cutoff, 44.1% of patients in the training cohort had survived at least two years. The best models achieved AUROC values between 0.63 and 0.74 for survival prediction - somewhat lower than CR prediction, reflecting the greater complexity of long-term outcomes.
The 25 features selected for survival prediction overlapped substantially with those for CR. Again, patient age was the most important single predictor - each additional year was associated with a 4.3% reduction in the odds of surviving two years. Hemoglobin level at diagnosis was also significant - for each 1 mmol/L increase up to normal levels, two-year survival odds rose by 14%.
Favorable genetic markers for survival included t(8;21), inv(16), NPM1 mutations, double-mutated CEBPA, and CEBPA-bZIP domain mutations. Notably, bZIP mutations were an independently favorable marker even when accounting for biallelic CEBPA status - confirming emerging evidence of their distinct prognostic significance.
Markers associated with worse two-year survival included mutations in TP53, DNMT3A, U2AF1, SF3B1, and high-level FLT3-ITD, as well as higher white blood cell counts, more circulating blast cells, elevated lactate dehydrogenase (a marker of cell turnover), and chromosomal deletions del(5) and del(17).
A critical test of any predictive model is whether it performs well on patients it has never seen before. The researchers validated their trained models on 664 AML patients from the AMLCG trials - a completely independent dataset not used in training or internal testing.
In external validation, CR prediction maintained strong performance with AUROC values between 0.71 and 0.80. Two-year overall survival prediction also held up, with AUROC between 0.65 and 0.75. These results are comparable to the internal test set, demonstrating that the models generalize well to new patient populations.
Notably, two prognostic variables - FLT3-TKD mutations and IKZF1 status - were unavailable in the validation cohort. Despite this missing information, the models still performed robustly. This suggests the overall feature architecture is redundant enough that losing a few inputs does not collapse accuracy.
The best-performing algorithm shifted between internal and external validation: random forest topped internal testing for CR prediction, while RBF-SVM (radial basis function support vector machine) was superior in external validation. This is exactly why the researchers built a pipeline with nine algorithms rather than one - different methods suit different datasets.
A key strength of the ML approach in this study is that it confirmed established risk markers while also shedding light on markers of previously uncertain relevance. The algorithms selected genes and chromosomal changes already recognized by clinical guidelines - such as TP53, ASXL1, NPM1, FLT3-ITD, and biallelic CEBPA - validating their importance independently of human assumptions.
Beyond established markers, the study provided additional evidence for U2AF1 and SF3B1 mutations as independent adverse prognostic factors. These genes regulate RNA splicing - how genetic messages are processed in the cell - and are more commonly associated with myelodysplastic syndromes and secondary AML. Their independent predictive value in intensive-therapy AML was less well-established before this study.
CEBPA-bZIP domain mutations emerged as an independent favorable predictor for both CR and two-year survival, even after adjusting for double-mutated CEBPA status. This aligns with recent reports that bZIP mutations represent a particularly favorable subset within the CEBPA-mutated AML group and may warrant recognition as a distinct disease category.
The study also confirmed that increasing hemoglobin at diagnosis progressively improves odds of both CR and survival - an often overlooked clinical variable. This may reflect overall disease burden and patient fitness, and is a readily available measurement in routine clinical practice.
Prior studies using conventional statistical methods for CR prediction in AML reported AUROC values between 0.71 and 0.78 in very large cohorts of 2,000 to 4,500 patients. This study's ML models matched or exceeded those accuracies in a cohort of 1,383 patients - suggesting ML extracts more signal from available data than traditional approaches.
The advantage of ML lies in its ability to model non-linear relationships between variables, handle interactions between dozens of features simultaneously, and adapt to data imbalances (like the fact that more patients achieved CR than did not). Traditional logistic regression, while interpretable, makes strong assumptions about the shape of relationships in data that may not hold in complex biology.
The researchers built their pipeline to be scalable and reusable - most steps are automated so the framework can be adapted to other cancer types or other prediction tasks with minimal modification. This is an important step toward practical clinical implementation, where models need to be updated as new patient data accumulate.
A key design choice was incorporating both clinical variables and genetic data. A previous analysis showed that about one-third of variation in AML survival comes from demographic and clinical features - not genetics alone. Models relying solely on next-generation sequencing data miss this dimension, limiting their usefulness in settings where advanced genomic testing is unavailable or routine.
The practical goal of predicting CR and survival at initial diagnosis is to support individualized treatment decisions. Patients identified as likely to fail intensive induction chemotherapy could be directed toward less toxic alternatives, such as the combination of venetoclax and azacitidine, which has proven effective for older or less fit AML patients and may be better suited for those unlikely to tolerate or benefit from intensive regimens.
For patients predicted to achieve remission, the models could also guide decisions about post-remission therapy - particularly the timing and indication for allogeneic stem cell transplantation. Accurate risk stratification at diagnosis could help determine which patients truly need transplantation and which might be safely managed with chemotherapy alone.
The researchers acknowledge limitations: the study is retrospective, meaning models were trained on historical data where treatment decisions were not guided by ML predictions. All patients received conventional chemotherapy regimens, limiting applicability to modern targeted therapies. Future work will prospectively test these models and extend them to novel agents like midostaurin, gilteritinib, and other targeted drugs.
For broad clinical adoption, the authors argue that ML models must use commonly available data - not rare or expensive tests - and must be cost-effective, accurate, and robust. Their approach uses standard next-generation sequencing panels and routine clinical labs, placing it within reach of most academic medical centers treating AML patients.
This study demonstrates that supervised machine learning can accurately predict complete remission and two-year survival in AML using multimodal patient data available at the time of diagnosis. The approach is validated in a large external cohort and outperforms conventional statistical tools in predictive power.
By simultaneously analyzing clinical, laboratory, cytogenetic, and molecular genetic data, ML models capture the full complexity of AML in a way that no single risk score can. The automated feature selection process allows the models to identify which factors matter - without a researcher having to decide in advance what to look for.
The authors propose their framework as a decision support system for hematologists - not a replacement for clinical judgment, but a tool that synthesizes complex multi-dimensional patient data into actionable risk estimates. The next step is prospective validation in real-time clinical settings, including in patients receiving newer targeted therapies.
Ultimately, as data from more patients accumulate and model performance improves with larger training sets, ML-guided risk assessment may become a standard part of AML diagnosis workflows - helping clinicians tailor treatment to the individual patient from day one.