Dermatofibrosarcoma protuberans (DFSP) is a rare cutaneous soft tissue sarcoma classified as having intermediate malignancy. Its incidence ranges from roughly 0.8 to 5 cases per million people per year, placing it in the category of truly rare tumors. DFSP arises most commonly in adults in their 20s through 40s, predominantly on the trunk, with a male-to-female ratio close to 1:1. While the local recurrence rate after inadequate surgery is high, the metastasis rate for conventional DFSP is low, estimated at only 2 to 5%.
The fibrosarcomatous subtype: A subset of DFSP cases undergo a process called fibrosarcomatous (FS) transformation, producing a more aggressive tumor variant known as fibrosarcomatous DFSP (FS-DFSP). Penner first described this phenomenon in 1951. Across multiple histopathological series, FS transformation is estimated to occur in 5 to 15% of all DFSP cases. FS-DFSP carries a substantially higher risk of local recurrence and distant metastasis compared to conventional DFSP (C-DFSP), and distinguishing the two subtypes at diagnosis is clinically essential for guiding treatment intensity and surgical margins.
The diagnostic gap: Despite its clinical significance, relatively few large-scale studies had systematically compared the clinicopathological features of FS-DFSP and C-DFSP. Distinguishing these variants requires attention to histological growth patterns, immunohistochemical staining profiles, and clinical behavior. This uncertainty creates risk of misclassification and, consequently, undertreated disease. The authors at West China Hospital of Sichuan University conducted this study to fill that gap, analyzing 221 DFSP patients and adding an AI-based classification layer via a back-propagation (BP) neural network trained on clinical variables.
The BP neural network component is notable because it demonstrates that an artificial intelligence tool can be constructed from purely clinical and demographic inputs (sex, age, tumor location, growth pattern) rather than requiring expensive molecular testing or specialized imaging, making the approach potentially deployable in resource-limited settings.
This was a retrospective cohort study conducted at the West China Hospital of Sichuan University, covering patients diagnosed with DFSP between 2010 and 2019. All 221 patients were Chinese, and diagnosis in every case was based on histological data rather than clinical presentation alone. The Ethics Committee of West China Hospital approved the study, and written informed consent was obtained from all patients, or from parents when patients were minors. Histopathological review classified each case as either conventional DFSP (C-DFSP, n=195) or fibrosarcomatous DFSP (FS-DFSP, n=26), reflecting an FS-DFSP frequency of 11.7% in this cohort.
Variables collected: For every patient, the investigators recorded sex, age at presentation (when the patient first noticed the tumor), age at the time of formal diagnosis, the interval between presentation and diagnosis, tumor location, tumor size at presentation, tumor size at diagnosis, overall tumor growth, annual tumor growth, growth type (indolent, gradually increasing, or rapidly enlarging), history of prior trauma, presence of pain, metastasis, recurrence, and follow-up duration.
Statistical methods: Categorical variables were analyzed with the Pearson chi-square test or Fisher's exact test. Continuous variables were compared using Student's t-test. All analyses were performed in SPSS 25.0 for Windows. A P value below 0.05 was considered statistically significant. Means were reported for continuous variables, and proportions for categorical ones. Immunohistochemistry (IHC) data, including CD34, CD10, SMA, desmin, S100, P53, EMA, and Ki-67 staining results, were also systematically collected for comparison between groups.
The BP neural network model was constructed separately using MATLAB's neural network toolbox with the Levenberg-Marquardt optimization algorithm. The model was designed with 10 input nodes (corresponding to the clinical variables), 3 hidden neurons, and 1 output node representing the binary classification target: 1 for FS-DFSP, 0 for C-DFSP. The model was trained on labeled clinical data and then tested on a held-out test set to assess generalization performance.
The most important finding in the clinical feature comparison is that FS-DFSP and C-DFSP are strikingly similar across most standard demographic and dimensional metrics. Age at presentation (mean 35.4 years for FS-DFSP vs. 30.7 for C-DFSP, P=0.091), age at diagnosis (40.5 vs. 37.3 years, P=0.168), sex distribution (male-to-female 18:8 in FS-DFSP vs. 108:87 in C-DFSP, P=0.18), tumor size at presentation (1.0 cm vs. 1.1 cm, P=0.818), tumor size at diagnosis (3.1 cm vs. 2.7 cm, P=0.362), and the interval between presentation and diagnosis (5.1 vs. 6.5 years, P=0.302) were all non-significant. This clinical similarity is precisely what makes FS-DFSP dangerous: it can be indistinguishable from the less aggressive variant at initial evaluation.
Growth type is the key clinical discriminator: While size and demographics were similar, growth behavior differed dramatically. In FS-DFSP, 80.8% of tumors showed rapid enlargement, compared to only 38.9% of C-DFSP tumors (P less than 0.001). Conversely, only 3.8% of FS-DFSP cases were indolent, compared to 14.9% of C-DFSP cases. The time from tumor discovery to rapid enlargement was also significantly shorter in FS-DFSP: median 0.6 years (range 0 to 5.0) versus 4.0 years in C-DFSP (P=0.003). This suggests that rapid clinical acceleration of a previously stable lesion is a warning sign for fibrosarcomatous transformation.
Location and anatomical distribution: C-DFSP occurred most frequently on the chest (26.2%) and abdomen (25.1%). FS-DFSP showed a shift in anatomical preference, occurring most often on the chest (30.8%) and posterior thighs (30.8%), with no cases recorded on the upper or lower extremities. While not statistically significant (P=0.272), this distribution difference may reflect underlying biological differences.
Recurrence rates also differed substantially: 7 of 26 FS-DFSP cases (26.9%) recurred during the follow-up period (mean 4.9 years), compared to 4 of 195 C-DFSP cases (2.1%). One FS-DFSP case developed lung metastasis, while no C-DFSP cases metastasized (P=0.118, non-significant due to small event numbers). The overall predominance of male patients (male-to-female ratio of 1:0.75) in this series contrasts with some prior studies reporting near-equal or female-predominant distributions.
Immunohistochemistry is the principal laboratory tool for separating DFSP from other spindle cell tumors and for identifying the fibrosarcomatous subtype. In this study, several IHC markers were systematically measured across both groups, and two emerged with statistically significant differences: CD34 and Ki-67.
CD34 negativity in FS-DFSP: CD34 is a cell-surface glycoprotein widely used as a positive marker for DFSP. In conventional DFSP, CD34 staining is typically diffuse and strong, with reported positivity rates of 92 to 100%. In this cohort, CD34 was positive in 99.5% of C-DFSP cases (194 of 195), consistent with prior literature. However, in FS-DFSP, CD34 positivity dropped to 88.5% (23 of 26), making the CD34-negative rate 11.5% versus 0.5% in C-DFSP (P=0.005). Some prior studies have reported CD34 negativity in up to 50% of FS-DFSP cases, suggesting this cohort may represent a milder phenotype or reflect sampling differences. Regardless, significant CD34 loss in a DFSP-like lesion should raise suspicion for fibrosarcomatous transformation.
Ki-67 proliferative index: Ki-67 is a nuclear protein that marks actively proliferating cells. Its index (the percentage of tumor cells staining positive) is a direct measure of tumor proliferative activity. In this cohort, the mean Ki-67 index was 8.1% (SD 4.7%) for C-DFSP versus 18.1% (SD 12.2%) for FS-DFSP (P less than 0.001). This more than twofold difference in proliferative activity aligns with prior literature: Sasaki et al. reported values of 8.9% for C-DFSP and 21.5% for FS-DFSP. High Ki-67 is clinically relevant beyond DFSP, as elevated indices have been linked to worse outcomes in prostate cancer and mantle cell lymphoma, further validating its role as a cross-entity proliferative biomarker.
Other IHC markers (CD10, SMA, desmin, S100, EMA) showed no significant differences between groups. P53 positivity was found in 2 FS-DFSP cases and no C-DFSP cases (P=0.015), consistent with prior reports of P53 overexpression contributing to fibrosarcomatous progression. Activation of the Akt/mTOR, STAT3, ERK, and PD-L1 pathways has also been implicated in DFSP progression based on prior molecular studies, suggesting that fibrosarcomatous transformation is driven by accumulated molecular alterations rather than a single mutational event.
The BP neural network was constructed as a practical clinical decision-support tool, designed to classify DFSP cases as either conventional or fibrosarcomatous using only routinely available clinical information. This approach is meaningful because it bypasses the need for specialized molecular assays and relies entirely on data that any dermatologist or surgeon could collect during a standard consultation.
Network architecture: The model uses a feedforward, multilayer topology with 10 input nodes, 3 hidden layer nodes, and 1 output node. The 10 input features are: sex (X0), age at presentation (X1), age at diagnosis (X2), the interval between presentation and diagnosis (X3), tumor location (X4), tumor size at presentation (X5), tumor size at diagnosis (X6), total tumor growth (X7), annual tumor growth (X8), and growth type (X9). The output node produces a score that maps to 1 (FS-DFSP) or 0 (C-DFSP). The number of hidden nodes was set at 3 after iterative tuning.
Training algorithm: The Levenberg-Marquardt (LM) algorithm was selected for optimization. The LM algorithm is a hybrid method that transitions between gradient descent (efficient far from a minimum) and Gauss-Newton approximation (faster convergence near a minimum). It is particularly suited to small-to-medium sized datasets, where its faster convergence and numerical stability outperform pure gradient descent approaches like stochastic gradient descent or Adam. The LM algorithm is noted to minimize nonlinearities (local minima) with the fastest convergence speed, averaging 30 iterations in this study. The training target mean square error (MSE) was set at 0.01, and the model converged in 31 training runs.
The back-propagation algorithm itself, first proposed by Paul Werbos in 1974 and popularized in the 1980s by Rumelhart and colleagues, operates by propagating prediction errors backward through the network, computing gradients, and iteratively adjusting the weights of each connection to minimize overall error. Despite the availability of more modern deep learning architectures, BP networks remain useful for small, tabular datasets where the number of features is limited and training data is scarce, as is the case in rare cancers like DFSP.
The BP neural network achieved 100% classification accuracy on the training sample set, with a misdiagnosis rate of 0 on training data. This result is expected for a well-fitted neural network on its own training data, but it does not by itself indicate that the model generalizes to unseen cases.
Test set performance: When evaluated on a held-out test set, the model achieved a correct classification rate of 88.64% and a misdiagnosis rate of 11.36%. The overall reported classification and misdiagnosis rates across the full cohort (training plus test) were 84.1% and 15.9%, respectively. These figures indicate that the model generalizes reasonably well beyond the training data, though the test set performance drop from 100% to 88.64% is meaningful and underscores the importance of evaluating on independent data.
The authors note that the classification accuracy and feasibility of the BP neural network model are high for FS-DFSP, and that it may serve as a method for clinical auxiliary identification. This framing positions the tool as a decision aid rather than a replacement for histopathological diagnosis. In a rare cancer context where histopathology remains the gold standard but may not always be immediately available or unambiguous, an AI tool that achieves 84 to 88% accuracy from clinical inputs alone represents a potentially useful adjunct.
Context for the performance numbers: An 84.1% overall correct classification rate translates to a misclassification of approximately 1 in 6 cases. In a clinical setting, the consequences of misclassifying FS-DFSP as C-DFSP (false negative for the aggressive subtype) include under-resection, inadequate surgical margins, and higher recurrence risk. The model's utility would need to be weighed against the baseline diagnostic error rate of clinicians without AI assistance, which was not formally reported in this study.
The clinical stakes behind accurate FS-DFSP identification are directly tied to treatment selection. For conventional DFSP, two surgical approaches are established: wide local excision (WLE), historically the gold standard, and Mohs micrographic surgery (MMS), which assesses 100% of the resection margin while conserving maximal normal tissue. Recurrence rates after WLE range from 0 to 41% in published series, reflecting wide variation in surgical margin adequacy. After MMS, recurrence rates drop to 0 to 6.7%, making MMS the preferred approach at specialized centers.
FS-DFSP requires more aggressive treatment: FS-DFSP carries a higher rate of local recurrence after all surgical approaches. In multicenter European data cited in this paper, FS-DFSP patients experienced recurrence more frequently than C-DFSP patients after WLE. The authors' own prior work showed that FS change was an independent prognostic factor for local recurrence in both univariable and multivariable analyses after MMS, establishing it as one of the few histological features that can identify a higher-risk DFSP patient even after margin-negative surgery.
The misdiagnosis risk: When DFSP is treated as a benign mass with simple excision (without margin assessment), local recurrence rates reach 26 to 60%. Some FS changes are found only in recurrent tumors rather than the primary lesion, suggesting that incomplete excision may act as a selection pressure driving fibrosarcomatous transformation in residual disease. This underscores why preoperative identification of FS-DFSP matters: it should prompt wider margins, more rigorous intraoperative margin assessment, and closer postoperative surveillance.
For the rare cases of metastatic DFSP, imatinib mesylate (targeting the COL1A1-PDGFB fusion gene that drives DFSP pathogenesis) is the established systemic therapy. The metastatic risk is substantially higher in FS-DFSP than in C-DFSP, making accurate subtype identification relevant not just for surgical planning but also for systemic therapy preparedness.
Retrospective design: The authors acknowledge that the retrospective nature of this study is its primary limitation. Retrospective data collection introduces the risk of incomplete records, variable documentation standards across the 10-year study period, and selection bias toward more severe or diagnostically complex cases at a tertiary referral center. The long diagnostic interval (median 4.0 years from presentation to diagnosis in C-DFSP) reflects this referral bias, as community-diagnosed DFSP is likely diagnosed more quickly and is not fully represented.
Limited follow-up data: Long-term follow-up data were lacking for a substantial proportion of patients. The mean follow-up was 4.8 years (SD 2.2), which may be insufficient to capture late recurrences, particularly in FS-DFSP, where recurrences can occur years after apparently successful resection. The absence of comprehensive outcome data limits assessment of how well the BP neural network's classifications correlate with actual clinical prognosis over extended periods.
Dataset size and class imbalance: With only 26 FS-DFSP cases compared to 195 C-DFSP cases, the dataset is small and substantially imbalanced. Class imbalance is a recognized challenge in machine learning: models trained on imbalanced datasets tend to develop a bias toward the majority class. In the clinical context, this means the model may be more prone to classifying ambiguous cases as C-DFSP (the larger class), which is precisely the misclassification type with the highest clinical cost. Future iterations of this model would benefit from oversampling techniques (such as SMOTE), cost-sensitive learning, or larger multicenter datasets.
Future directions: The authors frame this work as a proof-of-concept for applying neural networks to rare cutaneous sarcoma classification. Expanding the dataset through multicenter collaboration across Chinese dermatology centers would be the most impactful next step. Incorporating IHC data (Ki-67 index, CD34 status) as additional input features could substantially improve model accuracy, since Ki-67 and CD34 are the two variables that differ most significantly between groups. External validation in a geographically distinct cohort, and eventual prospective evaluation in a clinical setting, would be required before the BP model could be considered for routine clinical use.