Lipomatous tumors are the most common soft tissue masses in clinical practice, driven largely by the high incidence of benign lipomas. However, within the malignant spectrum of soft tissue tumors, liposarcoma is one of the most frequently encountered subtypes. Well differentiated liposarcoma (WDLPS) is the largest subgroup: a low-grade, locally aggressive tumor defined at the molecular level by amplification of the MDM2 gene. In rare cases, WDLPS can transform into dedifferentiated liposarcoma (DDLPS), a far more aggressive histological subtype with a substantially worse prognosis.
The clinical challenge is that WDLPS and ordinary lipoma can look strikingly similar on routine MRI. Several imaging features have been described as distinguishing criteria, including tumor size, depth, location, and internal heterogeneity, but overlap between the two entities is considerable. Even experienced musculoskeletal radiologists frequently cannot separate them with confidence. This diagnostic ambiguity has real treatment consequences: lipomas generally require no surgery, while WDLPS patients are typically referred for resection and specialist oncological follow-up at a dedicated sarcoma center.
The current gold standard for diagnosis is an invasive tissue biopsy, which is then tested for MDM2 gene amplification using fluorescence in situ hybridization (FISH). MDM2 amplification is present in WDLPS but absent in lipoma. The biopsy procedure is painful, carries risks related to tumor location and proximity to neurovascular structures, and is subject to sampling error if the needle misses a representative region of the tumor. This study was designed to test whether radiomics, the extraction of quantitative features from routine imaging, could replace or reduce reliance on biopsy for this specific diagnostic decision.
Importantly, the authors restricted their model to the hardest clinical scenario: distinguishing WDLPS from lipoma, which requires the most nuanced differentiation. DDLPS, the aggressive subtype, has sufficiently distinct radiological features that experienced radiologists can identify it by eye, so it was included as a visual classification task for the radiologists but not incorporated into the radiomics model. This design choice made the radiomics task more clinically focused and more challenging than prior studies.
The retrospective cohort was drawn from all patients referred to or diagnosed at the Erasmus MC Cancer Institute in Rotterdam between December 2009 and August 2018. Eligibility required a pathologically confirmed diagnosis of lipoma, WDLPS, or DDLPS, a known MDM2 FISH status, and at least one pretreatment T1-weighted MRI scan. Because Erasmus MC is a regional sarcoma referral center, many MRI scans were obtained at outside hospitals rather than centrally, which introduced significant imaging heterogeneity into the dataset.
Cohort composition: In total, 138 tumors were included: 58 MDM2-negative lipomas, 58 MDM2-positive WDLPS, and 22 MDM2-positive DDLPS. The 116 lipoma and WDLPS scans, which formed the core dataset for radiomics model development, came from 41 different MRI scanners. Field strength varied across 1 T (10 scans), 1.5 T (98 scans), and 3 T (8 scans). Three manufacturers were represented: Siemens (45 scans), Philips (45 scans), and GE (26 scans) across 19 different scanner models. Slice thickness, repetition time, and echo time also varied widely.
Clinical characteristics: Most patients were male (60.1%), with tumors predominantly located in the lower extremity (51.4%) and trunk (26.8%). Median WDLPS size was 20.4 cm (interquartile range 15.9-26.3 cm) with a median volume of 36.3 cl (IQR 22.9-85.5 cl), compared with 12.3 cm (IQR 9.3-15.5 cm) and 12.9 cl (IQR 4.6-25.0 cl) for lipoma. This size difference turned out to be an important confounder. The vast majority of patients underwent surgery: 32 lipoma patients, 50 WDLPS patients, and 19 DDLPS patients.
Sequence availability: Beyond the mandatory T1 sequence, additional MRI sequences were present in subsets: T1 with fat saturation (T1-FS) in 55 patients (47.4%), T1 with gadolinium contrast (T1-GD) in 42 (36.2%), T1-FS with gadolinium (T1-FS-GD) in 80 (69.0%), T2 in 76 (65.5%), and T2 with fat saturation (T2-FS) in 92 (79.3%). This variability in sequence availability was handled by imputing missing feature values, rather than restricting the cohort to only patients with complete imaging sets.
The radiomics pipeline followed a well-defined sequence: manual segmentation of tumor regions, automated image registration across sequences, quantitative feature extraction using PyRadiomics, and machine learning classification using the WORC (Workflow for Optimal Radiomics Classification) toolbox. Each step was designed to handle the real-world messiness of multi-scanner, multi-protocol data.
Segmentation: Lipoma and WDLPS lesions were segmented semi-automatically on the T1 images to define regions of interest (ROIs). Segmentation was performed by either a medical master's student or a physician-researcher with an MD degree, both blinded to the tumor diagnosis. A musculoskeletal radiologist specialized in soft tissue sarcomas verified a sample set for quality control. For patients with additional MRI sequences, the T1 segmentation was transferred to the other sequences using automated image registration via the elastix software package, correcting for patient movement between scans.
Feature extraction: Using PyRadiomics, three categories of quantitative features were extracted from each ROI: intensity features (first-order statistics such as mean and standard deviation of pixel values), shape features (morphological properties including volume, surface area, and sphericity), and texture features (capturing heterogeneity, speckle patterns, and spatial relationships via gray-level co-occurrence matrices and similar approaches). Patient-level and manually scored features, including age, sex, tumor location, depth, lobularity, and atypical T1 appearance, were also assembled for comparison models.
WORC and ensemble modeling: Rather than manually selecting a single machine learning approach, WORC performed an exhaustive automated search across 100,000 candidate workflows, where each workflow represented a different combination of feature selection methods and classifiers. The 50 best-performing workflows were then combined into a single ensemble model. This ensemble approach was chosen to reduce the risk that any single best-performing solution was a coincidental finding, and to produce a more robust and generalizable result. Model evaluation used a 100-times random-split cross-validation, with 80% of data used for training and 20% for testing in each iteration, stratified to maintain the lipoma-to-WDLPS ratio.
The primary radiomics model, trained on T1 imaging features alone (model 1), achieved a mean AUC of 0.83 (95% CI 0.75-0.90), with a mean sensitivity of 0.68 and specificity of 0.84 across the 100 cross-validation iterations. In comparison, three expert radiologists with 3, 5, and 10 years of experience in soft tissue tumor imaging achieved AUCs of 0.74, 0.61, and 0.72 respectively. All three radiologist AUCs fell below the lower bound of the 95% confidence interval of the T1 imaging model, confirming that the model significantly outperformed human expert reading.
Where radiologists underperformed: The radiologists achieved sensitivity values comparable to (0.74 and 0.64) or higher (0.91) than the radiomics model, but their specificity was substantially lower. The T1 radiomics model achieved specificity of 0.84, while the three radiologists reached only 0.55, 0.36, and 0.59. This means that when radiologists correctly identified WDLPS, they did so partly by over-calling WDLPS, accepting a high rate of false positives on lipoma cases. The radiomics model was more selective. Interobserver agreement among the radiologists was poor, with Cohen's kappa values of 0.24, 0.04, and 0.40 between pairs, and a mean kappa of 0.23, a level that reflects only slight agreement beyond chance.
Adding T2 improves performance substantially: Most additional MRI sequences beyond T1 did not meaningfully improve the radiomics model. However, combining T1 and T2 imaging features produced a clear gain: the T1+T2 model achieved AUC 0.89 (95% CI 0.83-0.95), sensitivity 0.74, and specificity 0.88. All three radiologist AUCs also fell below the lower bound of this model's confidence interval. The authors confirmed the T2 benefit was genuine and not a result of selection bias in which patients happened to have T2 scans, as the distributions of tumor characteristics were comparable between patients with and without T2 available.
Comparison of other models: The model trained on patient features alone (age, sex, location) achieved AUC 0.75 with higher sensitivity (0.77) but lower specificity (0.59). Manually scored features (depth, lobularity, atypical T1 appearance) reached AUC 0.72, sensitivity 0.76, specificity 0.57. Combining imaging and manually scored features performed no better than imaging alone, suggesting radiomics features capture sufficient information and that observer-dependent manual features add noise rather than signal.
One of the most methodologically important aspects of this study is its analysis of the role of tumor volume. WDLPS tumors are on average considerably larger than lipomas, and this size difference risks becoming the dominant predictive signal in any machine learning model, masking whether true texture or intensity features carry independent diagnostic value. The authors specifically built model 5, a volume-only model, to measure this effect. Model 5 achieved an AUC of 0.83, sensitivity 0.67, and specificity 0.84, essentially matching the performance of the full T1 imaging radiomics model (model 1).
The volume distribution extremes: The authors found that 17 tumors with volume above 70 cl were all WDLPS, and 21 tumors with volume below 7 cl were all lipoma. These 38 cases represent a clinically easy subset where volume alone is sufficient. The interesting diagnostic question sits in the middle: 78 tumors with volumes between 7 and 70 cl, where size does not clearly differentiate the two entities. The authors created a volume-matched cohort from these cases, with similar volume distributions for lipoma and WDLPS, to stress-test whether radiomics features beyond volume provided genuine added value.
Volume-matched cohort results: On the volume-matched cohort, the T1 imaging model (model 1) dropped to AUC 0.69 (95% CI 0.58-0.80), sensitivity 0.60, specificity 0.74. The volume-only model dropped further to AUC 0.64. Crucially, the T1+T2 imaging model held up considerably better, achieving AUC 0.81 (95% CI 0.72-0.90), sensitivity 0.66, specificity 0.84 on this harder subset. The T1+T2 model substantially outperformed the T1 model (AUC 0.81 vs. 0.69) and the volume-only model (AUC 0.64) in this volume-matched setting, confirming that T2 texture and intensity features capture diagnostic information independent of tumor size.
Implications for future models: The finding that the full-cohort T1 model performs similarly to volume alone reveals a dataset limitation: the cohort contains few large lipomas or small WDLPS, biasing the model toward learning size as a proxy. The authors explicitly recommend expanding future datasets to include lipomas larger than 70 cl and WDLPS smaller than 7 cl, which would force the model to learn true imaging phenotype rather than size. Feature importance analysis on the volume-matched cohort identified 16 individually significant features after Bonferroni correction: 11 shape features (including volume-related statistics), 4 texture features, and 1 intensity feature.
The three musculoskeletal radiologists participating in this study had 3, 10, and 5 years of subspecialty experience respectively. Each was given access to all available MRI sequences for each patient, plus patient age and sex, before making their classification decisions using a 10-point confidence scale. The radiomics model, by contrast, used only MRI pixel data with no clinical covariates in its primary configuration, placing it at a modest informational disadvantage.
Full cohort performance: On the full 116-case lipoma/WDLPS cohort, AUCs for the three radiologists were 0.74, 0.72, and 0.61. All three fell below the lower bound of the T1 imaging model's 95% CI (0.75-0.90). Sensitivity values ranged from 0.64 to 0.91 across radiologists, but at the expense of very low specificity (0.36-0.59). Radiologist 2 achieved the highest sensitivity (0.91) but the lowest specificity (0.36), indicating a strong tendency to over-diagnose WDLPS rather than miss it. The objective, consistent nature of the radiomics predictions contrasts sharply with this variability.
Volume-matched cohort comparison: When restricted to the volume-matched subset, the radiologists achieved AUCs of 0.68, 0.74, and 0.55. The T1 imaging model matched this range (AUC 0.69), but the T1+T2 model (AUC 0.81, sensitivity 0.66, specificity 0.84) clearly outperformed all three readers in overall discrimination. Interobserver agreement on the volume-matched cohort was even worse than on the full cohort: Cohen's kappa values of 0.18, -0.04, and 0.34 between pairs, with a mean of 0.16. A kappa of -0.04 means the two radiologists agreed less than they would by random chance alone, a striking finding about inter-reader reliability in this diagnostic scenario.
DDLPS visual recognition: The same three radiologists also classified 22 DDLPS cases, to confirm that the more aggressive subtype can be identified visually without computational assistance. On this task, they achieved AUCs of 0.97, 0.91, and 0.90, with sensitivity values of 0.95, 0.95, and 0.91, and specificities of 0.95, 0.56, and 0.89. This confirms the clinical teaching that DDLPS is radiologically distinct from non-DDLPS lipomatous tumors, and validates the study's design decision to focus radiomics modeling on the genuinely ambiguous WDLPS/lipoma pair.
The authors discuss four specific limitations that must be understood before this model is interpreted as ready for clinical deployment. Each limitation points to a concrete methodological gap rather than a vague statement about needing more research.
Volume bias: As detailed in the results, the strong volume separation between WDLPS and lipoma in this dataset means that the full-cohort T1 model is likely partially learning tumor size as a surrogate for diagnosis. Future studies need to intentionally enrich the dataset with large lipomas (greater than 70 cl) and small WDLPS (less than 7 cl), categories that were absent in this cohort, to force the model to distinguish the two entities on genuine imaging phenotype rather than size alone.
Manual segmentation variability: All tumor ROIs were manually segmented, introducing interobserver and intraobserver variability. This variability propagates through the entire radiomics pipeline, affecting feature extraction and ultimately classification. Manual segmentation is also time-consuming and would not be practical at clinical scale. The authors acknowledge that automated segmentation tools would be needed for clinical implementation, but note these were not yet available for this tumor type at the time of the study.
Retrospective design and selection bias: The dataset was collected retrospectively, which risks selection bias. The lipoma subset in particular may be non-representative: most routine small lipomas are managed locally without referral to a sarcoma center, so the lipomas in this dataset skew toward large, atypical, or clinically uncertain cases. This actually made the dataset harder and more clinically relevant, but means the reported performance may not generalize to typical community-practice lipomas, which are easier to classify. Similarly, demographic distributions (e.g., no WDLPS under age 35, no lipomas in the head and neck) may not reflect the full clinical spectrum.
Protocol heterogeneity: The 41-scanner, multi-manufacturer setting was deliberately included to reflect real-world practice rather than controlled single-institution acquisition. While the authors frame this as a strength regarding generalizability, the lack of standardized acquisition parameters also introduces uncontrolled variation in feature values. The study argues that the model is robust to this variation by training and testing on heterogeneous data, but prospective external validation on entirely independent multi-institution cohorts would be required to fully substantiate this claim.
The authors make a measured but substantive argument for the clinical translational value of their approach. A radiomics model that can identify WDLPS non-invasively on routine MRI would spare patients the discomfort and risk of an imaging-guided biopsy, eliminate the substantial cost of the biopsy procedure and FISH molecular testing, and potentially enable earlier referral to specialist sarcoma centers where outcomes are better. Because the model uses only T1 (and optionally T2) sequences without gadolinium contrast, it requires no additional imaging beyond a standard pre-operative MRI protocol, making it immediately compatible with current diagnostic pathways.
Risk stratification rather than binary replacement: The authors do not propose replacing biopsy outright, but rather using the radiomics model to stratify patients. The clearest candidates for biopsy avoidance are: (1) patients at elevated procedural risk from biopsy due to tumor location near critical structures, and (2) patients where the radiomics model predicts MDM2 status with high confidence. For a WDLPS misclassified as lipoma, the authors note that short-term active surveillance appears to be a clinically safe management option in patients without symptoms or tumor growth, so a false negative in this direction would not immediately harm the patient if a structured follow-up protocol is in place.
Comparison with prior literature: Thornhill and colleagues published a similar MRI radiomics approach for lipoma-versus-liposarcoma discrimination using 44 cases and including other liposarcoma subtypes such as DDLPS and myxoid liposarcoma. The current study improves on that work in multiple respects: larger sample size (116 vs. 44 cases), restriction to the genuinely ambiguous WDLPS/lipoma pair only, FISH-confirmed ground truth on all cases (rather than relying on histology alone without molecular confirmation), requirement for contrast-free sequences only, and development on a multi-scanner heterogeneous dataset to improve generalizability.
Required next steps: The authors identify several priorities for follow-up research. The dataset should be expanded to include volume-extreme cases (large lipomas, small WDLPS) that were absent here. Future imaging protocols for studies in this space should require both T1 and T2 sequences, as T2 provided a substantial AUC gain (0.83 to 0.89 on full cohort) despite being available in only 65.5% of this cohort. Automated segmentation methods need to be developed and validated to make the pipeline practical at scale. Finally, prospective external validation in independent cohorts at other institutions is required before any clinical implementation could be responsibly recommended.