Clinicopathological Features of Fibrosarcomatous Dermatofibrosarcoma Protuberans and the Construction of a Back-Propagation Neural Network Recognition Model

Orphanet Journal of Rare Diseases 2021 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
DFSP and Its Dangerous Fibrosarcomatous Variant

Dermatofibrosarcoma protuberans (DFSP) is a rare cutaneous soft tissue sarcoma classified as having intermediate malignancy. Its incidence ranges from roughly 0.8 to 5 cases per million people per year, placing it in the category of truly rare tumors. DFSP arises most commonly in adults in their 20s through 40s, predominantly on the trunk, with a male-to-female ratio close to 1:1. While the local recurrence rate after inadequate surgery is high, the metastasis rate for conventional DFSP is low, estimated at only 2 to 5%.

The fibrosarcomatous subtype: A subset of DFSP cases undergo a process called fibrosarcomatous (FS) transformation, producing a more aggressive tumor variant known as fibrosarcomatous DFSP (FS-DFSP). Penner first described this phenomenon in 1951. Across multiple histopathological series, FS transformation is estimated to occur in 5 to 15% of all DFSP cases. FS-DFSP carries a substantially higher risk of local recurrence and distant metastasis compared to conventional DFSP (C-DFSP), and distinguishing the two subtypes at diagnosis is clinically essential for guiding treatment intensity and surgical margins.

The diagnostic gap: Despite its clinical significance, relatively few large-scale studies had systematically compared the clinicopathological features of FS-DFSP and C-DFSP. Distinguishing these variants requires attention to histological growth patterns, immunohistochemical staining profiles, and clinical behavior. This uncertainty creates risk of misclassification and, consequently, undertreated disease. The authors at West China Hospital of Sichuan University conducted this study to fill that gap, analyzing 221 DFSP patients and adding an AI-based classification layer via a back-propagation (BP) neural network trained on clinical variables.

The BP neural network component is notable because it demonstrates that an artificial intelligence tool can be constructed from purely clinical and demographic inputs (sex, age, tumor location, growth pattern) rather than requiring expensive molecular testing or specialized imaging, making the approach potentially deployable in resource-limited settings.

TL;DR: DFSP affects 0.8 to 5 per million people per year; its fibrosarcomatous subtype (FS-DFSP) occurs in 5 to 15% of cases and carries significantly higher metastasis and recurrence risk. This study analyzed 221 patients at a single Chinese academic center and built a BP neural network classification model using 10 clinical input variables.
Pages 2-3
Study Design, Patient Cohort, and Statistical Approach

This was a retrospective cohort study conducted at the West China Hospital of Sichuan University, covering patients diagnosed with DFSP between 2010 and 2019. All 221 patients were Chinese, and diagnosis in every case was based on histological data rather than clinical presentation alone. The Ethics Committee of West China Hospital approved the study, and written informed consent was obtained from all patients, or from parents when patients were minors. Histopathological review classified each case as either conventional DFSP (C-DFSP, n=195) or fibrosarcomatous DFSP (FS-DFSP, n=26), reflecting an FS-DFSP frequency of 11.7% in this cohort.

Variables collected: For every patient, the investigators recorded sex, age at presentation (when the patient first noticed the tumor), age at the time of formal diagnosis, the interval between presentation and diagnosis, tumor location, tumor size at presentation, tumor size at diagnosis, overall tumor growth, annual tumor growth, growth type (indolent, gradually increasing, or rapidly enlarging), history of prior trauma, presence of pain, metastasis, recurrence, and follow-up duration.

Statistical methods: Categorical variables were analyzed with the Pearson chi-square test or Fisher's exact test. Continuous variables were compared using Student's t-test. All analyses were performed in SPSS 25.0 for Windows. A P value below 0.05 was considered statistically significant. Means were reported for continuous variables, and proportions for categorical ones. Immunohistochemistry (IHC) data, including CD34, CD10, SMA, desmin, S100, P53, EMA, and Ki-67 staining results, were also systematically collected for comparison between groups.

The BP neural network model was constructed separately using MATLAB's neural network toolbox with the Levenberg-Marquardt optimization algorithm. The model was designed with 10 input nodes (corresponding to the clinical variables), 3 hidden neurons, and 1 output node representing the binary classification target: 1 for FS-DFSP, 0 for C-DFSP. The model was trained on labeled clinical data and then tested on a held-out test set to assess generalization performance.

TL;DR: Retrospective cohort, 221 patients (2010 to 2019), single Chinese academic center. FS-DFSP accounted for 11.7% of cases. Statistical comparisons used chi-square, Fisher's exact, and Student's t-test (SPSS 25.0). The BP neural network was built in MATLAB with 10 input clinical variables, 3 hidden nodes, and binary output (FS-DFSP vs. C-DFSP).
Pages 3-5
How FS-DFSP and C-DFSP Differ Clinically

The most important finding in the clinical feature comparison is that FS-DFSP and C-DFSP are strikingly similar across most standard demographic and dimensional metrics. Age at presentation (mean 35.4 years for FS-DFSP vs. 30.7 for C-DFSP, P=0.091), age at diagnosis (40.5 vs. 37.3 years, P=0.168), sex distribution (male-to-female 18:8 in FS-DFSP vs. 108:87 in C-DFSP, P=0.18), tumor size at presentation (1.0 cm vs. 1.1 cm, P=0.818), tumor size at diagnosis (3.1 cm vs. 2.7 cm, P=0.362), and the interval between presentation and diagnosis (5.1 vs. 6.5 years, P=0.302) were all non-significant. This clinical similarity is precisely what makes FS-DFSP dangerous: it can be indistinguishable from the less aggressive variant at initial evaluation.

Growth type is the key clinical discriminator: While size and demographics were similar, growth behavior differed dramatically. In FS-DFSP, 80.8% of tumors showed rapid enlargement, compared to only 38.9% of C-DFSP tumors (P less than 0.001). Conversely, only 3.8% of FS-DFSP cases were indolent, compared to 14.9% of C-DFSP cases. The time from tumor discovery to rapid enlargement was also significantly shorter in FS-DFSP: median 0.6 years (range 0 to 5.0) versus 4.0 years in C-DFSP (P=0.003). This suggests that rapid clinical acceleration of a previously stable lesion is a warning sign for fibrosarcomatous transformation.

Location and anatomical distribution: C-DFSP occurred most frequently on the chest (26.2%) and abdomen (25.1%). FS-DFSP showed a shift in anatomical preference, occurring most often on the chest (30.8%) and posterior thighs (30.8%), with no cases recorded on the upper or lower extremities. While not statistically significant (P=0.272), this distribution difference may reflect underlying biological differences.

Recurrence rates also differed substantially: 7 of 26 FS-DFSP cases (26.9%) recurred during the follow-up period (mean 4.9 years), compared to 4 of 195 C-DFSP cases (2.1%). One FS-DFSP case developed lung metastasis, while no C-DFSP cases metastasized (P=0.118, non-significant due to small event numbers). The overall predominance of male patients (male-to-female ratio of 1:0.75) in this series contrasts with some prior studies reporting near-equal or female-predominant distributions.

TL;DR: FS-DFSP and C-DFSP are clinically indistinguishable by size, age, or sex. The key discriminator is growth behavior: 80.8% of FS-DFSP tumors grew rapidly vs. 38.9% of C-DFSP (P less than 0.001), and median time to rapid enlargement was 0.6 years in FS-DFSP vs. 4.0 years in C-DFSP (P=0.003). Recurrence occurred in 26.9% of FS-DFSP vs. 2.1% of C-DFSP.
Pages 5-6
IHC Markers That Distinguish FS-DFSP from Conventional DFSP

Immunohistochemistry is the principal laboratory tool for separating DFSP from other spindle cell tumors and for identifying the fibrosarcomatous subtype. In this study, several IHC markers were systematically measured across both groups, and two emerged with statistically significant differences: CD34 and Ki-67.

CD34 negativity in FS-DFSP: CD34 is a cell-surface glycoprotein widely used as a positive marker for DFSP. In conventional DFSP, CD34 staining is typically diffuse and strong, with reported positivity rates of 92 to 100%. In this cohort, CD34 was positive in 99.5% of C-DFSP cases (194 of 195), consistent with prior literature. However, in FS-DFSP, CD34 positivity dropped to 88.5% (23 of 26), making the CD34-negative rate 11.5% versus 0.5% in C-DFSP (P=0.005). Some prior studies have reported CD34 negativity in up to 50% of FS-DFSP cases, suggesting this cohort may represent a milder phenotype or reflect sampling differences. Regardless, significant CD34 loss in a DFSP-like lesion should raise suspicion for fibrosarcomatous transformation.

Ki-67 proliferative index: Ki-67 is a nuclear protein that marks actively proliferating cells. Its index (the percentage of tumor cells staining positive) is a direct measure of tumor proliferative activity. In this cohort, the mean Ki-67 index was 8.1% (SD 4.7%) for C-DFSP versus 18.1% (SD 12.2%) for FS-DFSP (P less than 0.001). This more than twofold difference in proliferative activity aligns with prior literature: Sasaki et al. reported values of 8.9% for C-DFSP and 21.5% for FS-DFSP. High Ki-67 is clinically relevant beyond DFSP, as elevated indices have been linked to worse outcomes in prostate cancer and mantle cell lymphoma, further validating its role as a cross-entity proliferative biomarker.

Other IHC markers (CD10, SMA, desmin, S100, EMA) showed no significant differences between groups. P53 positivity was found in 2 FS-DFSP cases and no C-DFSP cases (P=0.015), consistent with prior reports of P53 overexpression contributing to fibrosarcomatous progression. Activation of the Akt/mTOR, STAT3, ERK, and PD-L1 pathways has also been implicated in DFSP progression based on prior molecular studies, suggesting that fibrosarcomatous transformation is driven by accumulated molecular alterations rather than a single mutational event.

TL;DR: CD34 was negative in 11.5% of FS-DFSP vs. 0.5% of C-DFSP (P=0.005). Ki-67 index was 18.1% in FS-DFSP vs. 8.1% in C-DFSP (P less than 0.001), a more than twofold difference in proliferative activity. P53 positivity was found in 2 FS-DFSP cases (P=0.015). SMA, desmin, S100, and EMA were negative in both groups.
Pages 6-7
Building the Back-Propagation Neural Network for FS-DFSP Recognition

The BP neural network was constructed as a practical clinical decision-support tool, designed to classify DFSP cases as either conventional or fibrosarcomatous using only routinely available clinical information. This approach is meaningful because it bypasses the need for specialized molecular assays and relies entirely on data that any dermatologist or surgeon could collect during a standard consultation.

Network architecture: The model uses a feedforward, multilayer topology with 10 input nodes, 3 hidden layer nodes, and 1 output node. The 10 input features are: sex (X0), age at presentation (X1), age at diagnosis (X2), the interval between presentation and diagnosis (X3), tumor location (X4), tumor size at presentation (X5), tumor size at diagnosis (X6), total tumor growth (X7), annual tumor growth (X8), and growth type (X9). The output node produces a score that maps to 1 (FS-DFSP) or 0 (C-DFSP). The number of hidden nodes was set at 3 after iterative tuning.

Training algorithm: The Levenberg-Marquardt (LM) algorithm was selected for optimization. The LM algorithm is a hybrid method that transitions between gradient descent (efficient far from a minimum) and Gauss-Newton approximation (faster convergence near a minimum). It is particularly suited to small-to-medium sized datasets, where its faster convergence and numerical stability outperform pure gradient descent approaches like stochastic gradient descent or Adam. The LM algorithm is noted to minimize nonlinearities (local minima) with the fastest convergence speed, averaging 30 iterations in this study. The training target mean square error (MSE) was set at 0.01, and the model converged in 31 training runs.

The back-propagation algorithm itself, first proposed by Paul Werbos in 1974 and popularized in the 1980s by Rumelhart and colleagues, operates by propagating prediction errors backward through the network, computing gradients, and iteratively adjusting the weights of each connection to minimize overall error. Despite the availability of more modern deep learning architectures, BP networks remain useful for small, tabular datasets where the number of features is limited and training data is scarce, as is the case in rare cancers like DFSP.

TL;DR: The BP neural network uses 10 clinical inputs, 3 hidden nodes, and 1 binary output. The Levenberg-Marquardt optimization algorithm converged in 31 training runs to reach a target MSE of 0.01. The architecture was deliberately lightweight to remain applicable in settings without specialized laboratory infrastructure.
Pages 7-8
Classification Accuracy and Performance Metrics of the BP Model

The BP neural network achieved 100% classification accuracy on the training sample set, with a misdiagnosis rate of 0 on training data. This result is expected for a well-fitted neural network on its own training data, but it does not by itself indicate that the model generalizes to unseen cases.

Test set performance: When evaluated on a held-out test set, the model achieved a correct classification rate of 88.64% and a misdiagnosis rate of 11.36%. The overall reported classification and misdiagnosis rates across the full cohort (training plus test) were 84.1% and 15.9%, respectively. These figures indicate that the model generalizes reasonably well beyond the training data, though the test set performance drop from 100% to 88.64% is meaningful and underscores the importance of evaluating on independent data.

The authors note that the classification accuracy and feasibility of the BP neural network model are high for FS-DFSP, and that it may serve as a method for clinical auxiliary identification. This framing positions the tool as a decision aid rather than a replacement for histopathological diagnosis. In a rare cancer context where histopathology remains the gold standard but may not always be immediately available or unambiguous, an AI tool that achieves 84 to 88% accuracy from clinical inputs alone represents a potentially useful adjunct.

Context for the performance numbers: An 84.1% overall correct classification rate translates to a misclassification of approximately 1 in 6 cases. In a clinical setting, the consequences of misclassifying FS-DFSP as C-DFSP (false negative for the aggressive subtype) include under-resection, inadequate surgical margins, and higher recurrence risk. The model's utility would need to be weighed against the baseline diagnostic error rate of clinicians without AI assistance, which was not formally reported in this study.

TL;DR: Training accuracy was 100% (MSE 0.01 at 31 runs). Test set accuracy was 88.64% with an 11.36% misdiagnosis rate. Overall accuracy across training and test was 84.1%. The gap between training and test performance reflects the challenges of generalization in small, imbalanced datasets (26 FS-DFSP vs. 195 C-DFSP cases).
Pages 8-9
Surgical Management and Why Subtype Classification Drives Treatment Decisions

The clinical stakes behind accurate FS-DFSP identification are directly tied to treatment selection. For conventional DFSP, two surgical approaches are established: wide local excision (WLE), historically the gold standard, and Mohs micrographic surgery (MMS), which assesses 100% of the resection margin while conserving maximal normal tissue. Recurrence rates after WLE range from 0 to 41% in published series, reflecting wide variation in surgical margin adequacy. After MMS, recurrence rates drop to 0 to 6.7%, making MMS the preferred approach at specialized centers.

FS-DFSP requires more aggressive treatment: FS-DFSP carries a higher rate of local recurrence after all surgical approaches. In multicenter European data cited in this paper, FS-DFSP patients experienced recurrence more frequently than C-DFSP patients after WLE. The authors' own prior work showed that FS change was an independent prognostic factor for local recurrence in both univariable and multivariable analyses after MMS, establishing it as one of the few histological features that can identify a higher-risk DFSP patient even after margin-negative surgery.

The misdiagnosis risk: When DFSP is treated as a benign mass with simple excision (without margin assessment), local recurrence rates reach 26 to 60%. Some FS changes are found only in recurrent tumors rather than the primary lesion, suggesting that incomplete excision may act as a selection pressure driving fibrosarcomatous transformation in residual disease. This underscores why preoperative identification of FS-DFSP matters: it should prompt wider margins, more rigorous intraoperative margin assessment, and closer postoperative surveillance.

For the rare cases of metastatic DFSP, imatinib mesylate (targeting the COL1A1-PDGFB fusion gene that drives DFSP pathogenesis) is the established systemic therapy. The metastatic risk is substantially higher in FS-DFSP than in C-DFSP, making accurate subtype identification relevant not just for surgical planning but also for systemic therapy preparedness.

TL;DR: WLE recurrence rates for DFSP range from 0 to 41%; MMS reduces this to 0 to 6.7%. FS-DFSP carries higher recurrence risk even after margin-negative MMS and is an independent prognostic factor in multivariable analysis. Simple excision without margin control produces 26 to 60% recurrence. Imatinib targets metastatic DFSP via the COL1A1-PDGFB fusion.
Pages 9-10
Study Constraints and the Path Forward for AI-Assisted DFSP Diagnosis

Retrospective design: The authors acknowledge that the retrospective nature of this study is its primary limitation. Retrospective data collection introduces the risk of incomplete records, variable documentation standards across the 10-year study period, and selection bias toward more severe or diagnostically complex cases at a tertiary referral center. The long diagnostic interval (median 4.0 years from presentation to diagnosis in C-DFSP) reflects this referral bias, as community-diagnosed DFSP is likely diagnosed more quickly and is not fully represented.

Limited follow-up data: Long-term follow-up data were lacking for a substantial proportion of patients. The mean follow-up was 4.8 years (SD 2.2), which may be insufficient to capture late recurrences, particularly in FS-DFSP, where recurrences can occur years after apparently successful resection. The absence of comprehensive outcome data limits assessment of how well the BP neural network's classifications correlate with actual clinical prognosis over extended periods.

Dataset size and class imbalance: With only 26 FS-DFSP cases compared to 195 C-DFSP cases, the dataset is small and substantially imbalanced. Class imbalance is a recognized challenge in machine learning: models trained on imbalanced datasets tend to develop a bias toward the majority class. In the clinical context, this means the model may be more prone to classifying ambiguous cases as C-DFSP (the larger class), which is precisely the misclassification type with the highest clinical cost. Future iterations of this model would benefit from oversampling techniques (such as SMOTE), cost-sensitive learning, or larger multicenter datasets.

Future directions: The authors frame this work as a proof-of-concept for applying neural networks to rare cutaneous sarcoma classification. Expanding the dataset through multicenter collaboration across Chinese dermatology centers would be the most impactful next step. Incorporating IHC data (Ki-67 index, CD34 status) as additional input features could substantially improve model accuracy, since Ki-67 and CD34 are the two variables that differ most significantly between groups. External validation in a geographically distinct cohort, and eventual prospective evaluation in a clinical setting, would be required before the BP model could be considered for routine clinical use.

TL;DR: Key limitations include retrospective design, short median follow-up (4.8 years), and severe class imbalance (26 FS-DFSP vs. 195 C-DFSP). Future improvements should incorporate Ki-67 and CD34 as additional inputs, use class-balancing techniques, expand to multicenter datasets, and validate externally before clinical deployment.