Peripheral nerve sheath tumors (PNSTs) are soft-tissue tumors arising from peripheral nerves, spanning a wide spectrum from fully benign to highly aggressive. Among benign peripheral nerve sheath tumors (BPNSTs), schwannomas account for up to 80% of cases, with neurofibromas comprising 10-24% of the remainder. Malignant peripheral nerve sheath tumors (MPNSTs) are far rarer, representing only 2% of all soft-tissue sarcomas, yet they carry an outsized clinical threat: approximately 60% of MPNST patients develop metastases over time, and 40-70% experience local recurrence. The median 5-year survival for high-grade MPNST remains poor, underscoring the urgency of correct preoperative diagnosis.
The NF1 connection: About 23-51% of MPNSTs arise in patients with neurofibromatosis type 1 (NF1), a hereditary tumor predisposition syndrome. NF1 patients face an 8-13% lifetime risk of developing an MPNST, making it the leading cause of mortality in this population. Sporadic and radiation-induced MPNSTs account for the remaining cases. In NF1 patients who already carry numerous neurofibromas throughout their bodies, identifying which lesion has undergone malignant transformation is particularly challenging and clinically consequential.
Why conventional MRI falls short: MRI is the primary imaging modality for evaluating PNSTs, offering excellent soft-tissue contrast and anatomical detail. However, no reliably distinctive MRI feature has been identified that allows radiologists to definitively separate BPNSTs from MPNSTs. Studies have shown that even experienced radiologists achieve accuracy of only around 50% in this task, essentially equivalent to a coin flip. This diagnostic uncertainty drives the need for invasive biopsies, which carry risks of pain, nerve damage, and sampling error, particularly for heterogeneous tumors like neurofibromas undergoing focal dedifferentiation.
This 2024 study, published in Cancers and conducted at Erasmus Medical Center in Rotterdam, investigates whether radiomics combined with machine learning can provide a more reliable, non-invasive method for preoperative MPNST classification. The study also compares multiple MRI scan combinations, manual versus semi-automatic segmentation, and direct head-to-head performance against experienced musculoskeletal radiologists.
The study enrolled patients treated at Erasmus Medical Center between 2000 and 2019 who had a confirmed pathological diagnosis of either MPNST or BPNST (neurofibroma or schwannoma). Patients were included only if they had non-metastatic disease at diagnosis and had undergone MRI as part of standard care. Patients with BPNSTs also required either surgical treatment or at least one year of follow-up without malignant transformation, ensuring that the benign designation was confirmed over time rather than assumed. In total, 35 MPNSTs (32%) and 74 BPNSTs (68%) were included, reflecting the real-world rarity of the malignant subtype relative to benign lesions.
Baseline characteristics: The two groups were broadly matched for age and sex, with mean ages of 43 years (MPNST) and 44 years (BPNST) and a similar female predominance in both groups (57% and 70% female, respectively). NF1 diagnosis was present in 53% of MPNST patients and 52% of BPNST patients. Spontaneous pain was common in both groups (70% vs. 59%). Two clinically meaningful differences emerged: preoperative motor deficits were significantly more frequent in the MPNST group (33% vs. 12%, p = 0.02), and MPNSTs were substantially larger, with a mean tumor volume of 208 cm3 compared to 54 cm3 for BPNSTs (p = 0.01).
MRI heterogeneity: The dataset included six MRI sequence types: T1-weighted (T1w), T2-weighted (T2w), T1 with fat saturation and gadolinium contrast (T1w-FS-GD), T1 SPIR with gadolinium (T1w-SPIR-GD), T2 with fat saturation (T2w-FS), and T2 STIR. Critically, scans were acquired across 37 different scanners from 23 model types, spanning Siemens, Philips, and General Electric equipment. This substantial imaging heterogeneity, arising because many patients were referred from other hospitals with their existing scans, is an important real-world feature that distinguishes this dataset from more controlled single-institution studies.
Clinical features extracted from medical records included age at operation, sex, NF1 or Schwannomatosis diagnosis, presence of spontaneous pain, and presence of preoperative motor deficits. These five variables were evaluated both as a standalone predictive model and as an additive component to the imaging-based radiomics models.
The radiomics workflow began with tumor segmentation on MRI. Manual segmentation was performed once per lesion by one of two observers (I.A. or E.M.), with random allocation between the two raters. Segmentations were carried out on whichever MRI sequence showed the tumor most clearly. To transfer the region of interest (ROI) from the reference sequence to all other available sequences, automated image registration was applied using a rigid transformation model with mutual information as the metric, implemented in the Elastix software package. All registered segmentations were visually inspected, and cases with zero overlap due to misregistration were excluded.
Semi-automatic segmentation via InteractiveNet: To address the labor-intensive nature of manual segmentation, the study also evaluated InteractiveNet, a minimally interactive deep learning segmentation tool. InteractiveNet requires the user to provide only six extreme points per tumor (two in each of three directions), which guide a convolutional neural network to generate the full segmentation mask. This framework had previously been validated across 13 soft-tissue tumor types, but this study was the first to test it on MPNSTs and neurofibromas specifically. Each semi-automatic segmentation was rated on a four-point quality scale (Excellent, Sufficient, Insufficient, Incorrect) to enable analysis of how segmentation quality affects downstream model performance.
Feature extraction: For each lesion on each available MRI sequence, 564 radiomics features were extracted covering intensity statistics, shape descriptors, and texture measures. Extraction used the WORC toolbox (v3.6.3), which internally calls PREDICT (v3.1.17) and PyRadiomics (v3.0.1). Categorical clinical features were converted into numerical form for inclusion in multi-feature models.
The WORC automated machine learning pipeline: Rather than selecting a single classifier and manually tuning it, the Workflow for Optimal Radiomics Classification (WORC) algorithm performs an exhaustive automated search across multiple stages including feature imputation, feature scaling, feature selection, feature resampling, and classification. It evaluates 1,000 candidate workflow combinations and selects the 100 highest-performing decision models, which are then ensembled into a single final model. This ensemble approach reduces the risk that any single "best" model is a coincidental finding. Model evaluation used a 100-times random-split cross-validation with an 80/20 train-test split. Within each training fold, an internal five-times random-split cross-validation was performed for hyperparameter optimization, keeping all model tuning strictly within the training partition.
Among all radiomics models using manual segmentations, performance was surprisingly similar across the different MRI sequence combinations, with mean AUC values ranging from 0.66 to 0.71 and substantially overlapping confidence intervals. The best single-sequence model was the T1-weighted MRI model, which achieved a mean AUC of 0.71 [95% CI: 0.56, 0.86], accuracy of 0.75 [0.67, 0.84], sensitivity of 0.30 [0.10, 0.50], and specificity of 0.93 [0.84, 1.00]. The T2w model performed comparably with an AUC of 0.70 [0.56, 0.84]. Combining T1w and T2w sequences together did not improve performance (AUC 0.68), suggesting that the two sequences encode largely redundant diagnostic information in this dataset.
Adding contrast-enhanced sequences: Adding contrast-enhanced T1 (T1w-FS-GD or T1w-SPIR-GD) or fat-suppressed T2 sequences to the base models also failed to improve the AUC meaningfully. The combination of T1w, T2w, and T1w-FS-GD or T1w-SPIR-GD reached AUC 0.70, essentially identical to the T1w model alone. The most parsimonious interpretation is that shape-based features dominate predictive signal in this dataset regardless of which MRI sequence they are measured on, and that additional sequences provide minimal incremental discriminative information once shape has been captured.
The sensitivity problem: Across all radiomics models, sensitivity was notably low, consistently falling in the range of 0.25-0.33. This means the models frequently missed true MPNSTs, classifying malignant tumors as benign. High specificity (0.87-0.95) came at the cost of low sensitivity, a tradeoff directly attributable to the class imbalance in the dataset (35 MPNSTs vs. 74 BPNSTs) and the tendency of the models to default toward the more common benign class. Adjusting the classification threshold along the ROC curve can shift this balance, but the underlying discriminative ability as measured by AUC remains modest.
Adding clinical features: Combining T1w MRI radiomics with the five clinical features produced a slight improvement in mean AUC from 0.71 to 0.74 [0.60, 0.88]. The clinical-features-only model yielded AUC 0.60 and the volume-only model AUC 0.64, both substantially below the radiomics models. The integrated T1w plus clinical model maintained accuracy of 0.75, specificity of 0.92, but sensitivity of only 0.31, confirming that clinical data adds modest discriminative value but does not resolve the sensitivity limitation.
The semi-automatic T1w radiomics model using InteractiveNet achieved a mean AUC of 0.68 [0.56, 0.79], compared to 0.71 [0.56, 0.86] for the manually segmented T1w model. Accuracy was 0.70 vs. 0.75, sensitivity 0.23 vs. 0.30, and specificity 0.94 vs. 0.93. The performance gap is small in absolute terms, and the 95% confidence intervals of the two models overlap substantially, meaning the observed difference is not statistically definitive. However, the directional finding is consistent: semi-automatic segmentation produces slightly inferior but broadly comparable results relative to manual segmentation.
Segmentation quality analysis: When the semi-automatic T1w model was restricted to only lesions where InteractiveNet produced "Sufficient" or "Excellent" quality segmentations, performance remained essentially unchanged compared to the model using all segmentations, regardless of quality rating. This is an encouraging finding because it suggests that quality filtering does not rescue poor segmentations, and that the algorithm performs consistently even across the range of cases where the observer rated quality as acceptable. The implication is that the interactive segmentation approach provides a stable contribution to the pipeline across typical clinical cases.
Clinical relevance: Manual segmentation of soft-tissue tumors on MRI is a significant bottleneck in clinical radiomics pipelines. Each manual contour must be drawn slice-by-slice by a trained observer, verified, and then propagated through multiple sequence registrations. For rare tumors like MPNSTs, where prospective validation datasets need to be built iteratively over time, reducing per-case annotation burden is practically important. InteractiveNet's requirement for only six extreme point clicks per tumor dramatically reduces operator time, particularly for tumors located in less accessible or visually complex anatomical regions.
This study also confirms that InteractiveNet, previously validated on 13 other soft-tissue tumor types, generalizes reasonably to MPNSTs and neurofibromas, which had not been included in prior validation work. The modest performance sacrifice may be acceptable in the context of pilot studies or large-scale retrospective dataset construction, where full manual annotation of every case is impractical.
Two musculoskeletal radiologists with 8 and 9 years of experience in soft-tissue sarcoma evaluation independently scored each lesion on a four-point scale indicating confidence that the tumor was an MPNST. They rated all 108 cases (the full patient set) without clinical information first, then repeated the assessment with clinical features available. Radiologist 1 achieved an imaging-only AUC of 0.79 [0.68, 0.88], accuracy of 0.79, sensitivity of 0.71, and specificity of 0.82. Radiologist 2 performed considerably lower: imaging-only AUC 0.68 [0.58, 0.78], accuracy 0.64, sensitivity 0.60, specificity 0.66. Cohen's kappa between the two radiologists was 0.47 without and 0.46 with clinical features, indicating only moderate interobserver agreement.
The clinical features paradox: Counterintuitively, providing clinical features to the radiologists slightly decreased their mean AUCs (Radiologist 1: 0.79 to 0.75; Radiologist 2: 0.68 to 0.66). Analysis of changed decisions revealed that most errors introduced in the second round came from reclassifying benign tumors with above-average volume as malignant and leaving malignant tumors with below-average volume classified as benign. This finding suggests that volume dominated the radiologists' reasoning when clinical data was provided, overriding nuanced imaging-based assessments and in some cases leading them astray.
Integrated models combining radiomics and radiologist opinions: The study tested four integration strategies using logical OR and AND operations between the binarized T1w plus clinical model and each radiologist's binary predictions. OR integration (tumor flagged as malignant if either the model or radiologist predicts malignancy) improved sensitivity substantially compared to the radiologists alone (sensitivity 0.74 with Radiologist 1, 0.71 with Radiologist 2) but at the cost of lower specificity (0.65 and 0.61, respectively) and overall accuracy (0.68 and 0.63). AND integration (malignancy requires agreement from both model and radiologist) achieved high specificity (0.96 and 0.93) and accuracy (0.78 and 0.75) but drastically reduced sensitivity to 0.21 and 0.20.
Neither integration strategy improved overall AUC compared to Radiologist 1 operating alone. However, the OR approach offers a clinically useful safety net for settings where missing an MPNST carries a higher cost than investigating a false positive, while the AND approach may support watchful waiting decisions when both independent sources agree a lesion is benign.
Small and imbalanced dataset: With only 35 MPNSTs and 74 BPNSTs, the dataset is limited by the rarity of the disease. This class imbalance directly contributes to the low sensitivity observed across all radiomics models. Increasing the malignant case count would be essential for training a model with clinically meaningful sensitivity. The retrospective nature of the cohort also introduces selection bias: BPNSTs referred to a tertiary sarcoma center are more likely to be symptomatic or large, making them more similar to MPNSTs and thus harder to classify correctly than a random sample of community-diagnosed BPNSTs.
Imaging heterogeneity: The dataset originated from 37 different scanners across 23 model types from three manufacturers. This heterogeneity in acquisition parameters likely degraded radiomics feature reproducibility and contributed to the modest AUC values observed. Radiomic features, particularly texture-based descriptors, are known to be sensitive to differences in scanner type, field strength, slice thickness, and reconstruction kernel. Studies using standardized single-protocol acquisitions, such as Ristow et al., who reported AUC 0.94 using fat-suppressed T2w MRI in a single-institution standardized dataset, achieved substantially higher performance. Zhang et al. obtained AUC 0.85 using T1-GD scans across three institutions with a larger cohort of 95 MPNSTs and 171 BPNSTs. The present study's more modest results likely reflect the real-world challenge of multi-scanner data rather than a fundamental limitation of radiomics in this application.
Missing data and imputation: Because no standardized imaging protocol exists for PNST diagnosis across referring hospitals, not all patients had all six MRI sequence types available. WORC employed automated feature imputation to handle missing sequences, but imputation introduces uncertainty and may have diluted the discriminative signal from contrast-enhanced sequences in models that nominally included them.
Low sensitivity across all models: The most clinically serious limitation is that all radiomics models exhibited sensitivity values of 0.25-0.33, meaning they would miss 67-75% of actual MPNSTs at the operating point tested. For a condition where delayed diagnosis leads to unresected high-grade sarcoma, this level of sensitivity is insufficient as a standalone diagnostic replacement for biopsy. The models may still offer value as triage tools or in combination with radiologist assessment, but they cannot currently substitute for histopathological confirmation.
Additional imaging modalities: FDG-PET is already used as a supplementary modality for distinguishing plexiform neurofibromas from MPNSTs in NF1 patients, and multiple studies have explored SUVmax thresholds as a discriminative parameter. However, optimal cut-off values remain variable across institutions and patient populations. Diffusion-weighted imaging (DWI) offers another promising avenue: apparent diffusion coefficient (ADC) values are generally higher in BPNSTs than in MPNSTs, reflecting differences in cellularity. Integrating ADC-derived radiomic features, PET-derived metabolic features, and structural MRI radiomics into a combined multimodal model could substantially improve discriminative performance beyond what any single modality achieves alone.
Deep learning segmentation and classification: The paper's authors suggest that deep learning end-to-end models, which learn directly from raw image pixels rather than pre-specified handcrafted features, may eventually surpass traditional radiomics pipelines for MPNST classification. However, the rarity of MPNSTs poses a fundamental challenge for data-hungry deep learning architectures. Potential solutions include transfer learning from large soft-tissue sarcoma datasets to rare MPNST-specific fine-tuning tasks, federated learning across multiple sarcoma reference centers to pool cases without sharing raw patient data, and generative data augmentation using models like variational autoencoders or GANs to synthetically expand the malignant training set.
Prospective multi-institution studies: The most important future step is prospective validation in a multi-institutional setting with standardized acquisition protocols. Retrospective single-institution data, even when carefully curated, carries inherent biases that prospective studies can avoid. Multi-center consortia for rare sarcoma subtypes, modeled on existing initiatives such as the European Organisation for Research and Treatment of Cancer (EORTC) soft tissue sarcoma group, could provide the combined case volumes needed to train models with clinically meaningful sensitivity and to evaluate generalizability across diverse patient populations and imaging environments.
Clinical integration and decision support: Future work should also address how radiomics predictions can be practically integrated into clinical workflows, ideally as a quantitative adjunct to radiologist reporting rather than a replacement. Explainability tools such as gradient-weighted class activation mapping (Grad-CAM) or SHAP values could help clinicians understand which imaging features are driving predictions in individual cases, improving trust and enabling appropriate integration with clinical judgment. Defining the specific clinical scenario where the model adds the most value, for example flagging borderline cases for expert review or providing a second opinion in centers without musculoskeletal radiology subspecialty expertise, will be key to translating research findings into practice.