Axillary lymph node status is one of the strongest prognostic factors in breast cancer, and the standard method for assessing it is sentinel lymph node biopsy (SLNB) -- the surgical removal and examination of the first lymph nodes draining the tumor. While SLNB is safer than full lymph node dissection, it still carries risks including lymphedema, shoulder dysfunction, and arm morbidity.
A critical distinction exists between different sizes of lymph node involvement. Macrometastasis (a metastatic deposit larger than 2 mm) is clinically significant and typically triggers more aggressive treatment decisions including completion axillary dissection and post-mastectomy radiotherapy. Micrometastasis and isolated tumor cells, in contrast, are now considered of minor clinical significance and often do not change management. This means that identifying specifically which patients have macrometastasis is the most actionable prediction target.
The trend toward de-escalation of axillary surgery has accelerated following evidence that SLNB provides no survival benefit in many patient groups. Guidelines now recommend omitting SLNB in women over 70 with early-stage hormone receptor-positive tumors. Noninvasive tools that could identify patients with a low risk of sentinel lymph node macrometastasis (macro-SLNM) would allow broader and safer SLNB omission without compromising oncological outcomes.
This study from Lund University is the first to systematically evaluate deep learning (DL) models for predicting sentinel lymph node macrometastasis using exclusively preoperative clinicopathological variables -- data already available before surgery from standard diagnostic workup. Unlike most prior studies that used imaging modalities, this study worked with structured tabular data alone.
The dataset comprised 18,185 patients from the Swedish National Quality Registry for Breast Cancer (NKBC), a nationwide population-based register with high coverage and low missing data rates. Patients diagnosed between 2014 and 2016 formed the development set (n = 13,656), and those diagnosed in 2017 were the temporal test set (n = 4,529). This time-based split tests the model's ability to generalize across calendar years, reflecting how it would be deployed in a prospective clinical setting.
Five machine learning algorithms of increasing complexity were compared: logistic regression (LR), multilayer perceptron (MLP), ResNet, Transformer, and CatBoost. A univariable model based on tumor size alone served as an additional benchmark to assess how much value the full feature set actually added.
Thirteen clinicopathological features available preoperatively were used: age, menstrual status, mode of cancer detection (screening vs. symptomatic), number of invasive foci, tumor stage, tumor size, histological grade, histopathological type, ER, PgR, HER2, Ki67 expression, and surrogate molecular subtype. Notably, lymphovascular invasion (LVI) -- a known predictor of nodal spread -- was excluded because it cannot be reliably assessed in a preoperative biopsy setting.
The Transformer model is architecturally designed around self-attention mechanisms, originally developed for natural language processing. Applied to tabular data, it uses feature tokenization -- converting each clinical variable into a learnable embedding vector, analogous to word embeddings in language models -- before applying attention to capture relationships among features. This allows the model to detect complex non-linear interactions that logistic regression cannot represent.
To address the strong class imbalance (only 13% of patients had macro-SLNM), the models were trained with three advanced loss strategies: weighted binary cross-entropy (upweighting the minority class), focal loss (emphasizing hard-to-classify samples), and triplet loss (training the model to distinguish similar and dissimilar cases). Hyperparameters were optimized using Bayesian search via the Optuna library, with up to 100 settings evaluated per algorithm.
On the test set, all five multivariable models achieved similar ROC AUCs in the range of 0.704 to 0.712, with Transformer, LR, and MLP all reaching approximately 0.711-0.712. The univariable tumor-size-only model achieved 0.692 -- just 2 percentage points lower. This surprisingly small improvement from adding 12 additional clinical features suggests that most of the predictive information in the available variables is already captured by tumor size alone.
Precision-recall (PR) AUC, which is more sensitive to performance on the minority class (macro-SLNM positive patients), showed LR performing best at 0.273, compared to 0.267 for Transformer and 0.239 for tumor size alone. Again, the improvements over the univariable baseline were marginal. This pattern across all metrics consistently indicated that there is limited additional information in the 13-feature set beyond what tumor size already provides.
Advanced DL strategies including focal loss, triplet loss, and extensive hyperparameter optimization through Bayesian search did not meaningfully improve performance on the internal validation set. The only strategy that clearly helped was the feature tokenizer for Transformer, which was found to be essential for that architecture to function effectively on tabular data.
The study's most clinically meaningful finding emerged when models were evaluated under the constraint of achieving at least 90% sensitivity -- matching the accepted false-negative rate of actual SLNB (10%). At this threshold, Transformer achieved the best specificity of 34.6%, precision of 16.2%, negative predictive value (NPV) of 96.2%, and accuracy of 41.5% -- outperforming LR (specificity 32.6%, NPV 95.9%) and all other models.
The NPV of 96.2% means that when the Transformer model predicts absence of macro-SLNM, it is correct about 96 times out of 100 -- comparable to the benchmark set by SLNB itself. In a clinical context where the goal is to identify patients who can safely skip SLNB, this NPV is the most relevant metric: a high NPV means few true macro-SLNM cases would be missed.
The consistent superiority of DL models over LR specifically at the 90% sensitivity operating point, despite similar overall AUC values, illustrates a general property of these architectures: DL models may learn decision boundaries that are better calibrated at specific operating points even when their overall discriminative ability is comparable to simpler models. This makes the specific operating point analysis an important complement to AUC comparison.
SHAP (Shapley Additive exPlanations) analysis was used to quantify how much each feature contributed to each model's predictions. Tumor size was the single most important predictor in all five models by a substantial margin -- its SHAP importance was significantly higher than the second-ranked variable in every model (p less than 0.001). The number of invasive foci (multifocality) ranked second in three of the five models.
Histological grade, hormone receptor status (ER, PgR), HER2, Ki67, and molecular subtype contributed only minimally to prediction accuracy. This finding explains why the multivariable model barely outperforms the tumor-size-only model: the additional features add little incremental information once tumor size is known. The high correlation between Transformer's SHAP values and effect sizes (r = 0.965) confirmed that the model's feature weighting accurately reflected clinical reality.
Individual patient-level SHAP analysis revealed the fundamental data limitation: true-positive and false-positive predictions often had similar feature profiles (both showing large tumor size and multifocality), and true-negative and false-negative predictions were also similar to each other (small unifocal tumors). Patients with essentially identical clinical profiles had divergent SLNB outcomes, indicating that there are important biological factors driving nodal spread that are simply not measured by standard clinicopathological variables.
This study was motivated by a broader shift in breast cancer surgery toward de-escalation of axillary treatment. The SOUND trial demonstrated that omitting axillary surgery was non-inferior for 5-year distant disease-free survival in patients with T1 tumors and negative axillary ultrasound -- supporting that many patients receive no benefit from SLNB. The ongoing INSEMA and BOOG 2013-08 trials are extending this evidence to broader patient groups.
Current guidelines (ASCO, NCCN) now endorse omitting SLNB in women over 70 with early-stage, hormone receptor-positive, HER2-negative disease. A validated noninvasive prediction tool could extend this guidance to younger patients and other molecular subtypes, particularly in patients planning mastectomy where accurate preoperative nodal assessment is especially important for decisions about post-mastectomy radiotherapy and immediate breast reconstruction.
Patients with macro-SLNM who undergo mastectomy with immediate breast reconstruction (IBR) face a particularly complex situation: post-mastectomy radiotherapy (PMRT), which is typically indicated when macrometastasis is found, significantly increases the risk of reconstruction failure and complications. Identifying these patients preoperatively would allow better surgical planning, including counseling about reconstruction options and timing before the patient is on the operating table.
The study's core finding -- that advanced DL architectures provide only marginal improvement over logistic regression on these 13 features -- reflects a fundamental principle in machine learning: model complexity is only beneficial when the data contains exploitable non-linear structure. The near-equal performance of all five models suggests the clinicopathological features available in this dataset have limited non-linear interactions, making them poorly suited for DL's key strength.
This outcome is consistent with ongoing debates in the machine learning community about whether DL outperforms gradient-boosted decision trees on tabular data generally. The study supports the view that DL advantages on tabular data emerge primarily when there are many features with complex interdependencies -- a situation not present here. The authors frame this as a data uncertainty problem: the target variable (SLNM status) appears nearly random given these features alone, representing irreducible aleatoric uncertainty.
The path forward identified by the authors is to increase input dimensionality with novel information sources rather than to apply more sophisticated models to the same data. Promising additions include deep radiomics from ultrasound or mammography, digital pathology analysis of H&E slides from core-needle biopsies, and gene expression signatures. Studies combining clinical features with imaging radiomics have achieved AUCs between 0.82 and 0.94, demonstrating the potential of multimodal approaches over clinical data alone.
This study established that deep learning models, particularly Transformer, can outperform logistic regression for macro-SLNM prediction under the clinically relevant constraint of high sensitivity, achieving an NPV of 96.2% at 90% sensitivity. However, the overall discriminative ability of all models (AUC 0.704-0.712) was only marginally better than tumor size alone (0.692), highlighting the fundamental limits of prediction from standard preoperative clinical features.
Key limitations include the retrospective design, the use of postoperative final pathology data for some features (particularly tumor size and multifocality, which may differ from preoperative estimates), and potential regional generalizability issues since macro-SLNM prevalence varies across populations. The use of only 13 features from a single national registry also meant that variables like lymphovascular invasion -- a known strong predictor -- could not be included.
The authors conclude that additional predictors with novel biological information, rather than more sophisticated modeling of existing features, are the essential next step. Integration of imaging-derived features, histopathological analysis from biopsy slides, and genomic signatures represents the most promising direction for achieving the level of accuracy needed to guide individual clinical decisions about SLNB omission.