Bone and soft tissue sarcomas represent one of the most diagnostically demanding areas in oncology. Sarcomas are believed to share a mesenchymal cell origin, yet histological observation reveals extraordinary variation, even within a single subtype. Chondrosarcoma alone, defined by the presence of cartilage-expressing lineages, can be subdivided into at least six distinct subtypes: central conventional, periosteal, extraskeletal myxoid, mesenchymal, clear cell, and dedifferentiated chondrosarcoma. The WHO guidelines, which have become the gold standard, currently define over 100 types and mimics of sarcoma, with many additional subtypes nested within them.
Historical context: Clinical recognition of bone and soft tissue tumours stretches back to antiquity. Hippocrates recorded observations of these tumours between 460 and 375 BC, and practitioners in ancient Rome identified distinct types including lipomas. Microscopic examination became standard practice following the 1712 publication of Etmullerus, and from there bone and soft tissue sarcoma evolved into distinct fields of pathological research. For much of the 20th century, classification systems were based on the resemblance of sarcomas to normal tissues and cell lineages, a framework that proved increasingly inadequate as molecular evidence accumulated.
The classification update problem: Each WHO edition has introduced notable amendments, reflecting the irreconcilable tension between histological ambiguity and emerging molecular evidence. A recent example from the 2020 edition is the redefinition of undifferentiated round cell tumours as entities distinct from Ewing sarcoma, with BCOR and CIC-rearranged sarcomas now recognized as separate diseases. These constant revisions are not arbitrary, they follow genuine biological discoveries. But they also mean that AI classifiers trained on older datasets may be classifying against an outdated taxonomy.
This 2025 commentary by William CH Cross, published in the Journal of Pathology: Clinical Research, surveys the state of pan-sarcoma datasets and AI-based classifiers, anchored by a detailed discussion of the DKFZ Sarcoma Classifier and its UK validation study. The paper is a measured scientific assessment of both the genuine promise and the persistent structural barriers to AI-driven sarcoma diagnosis.
The recognition that mutational patterns can improve sarcoma diagnoses has driven the gradual incorporation of molecular tests into clinical workflows. Specific driver events have become diagnostically useful, including nucleotide aberrations such as TP53 in osteosarcoma and gene fusions such as the EWSR1::FLI1 translocation that defines Ewing sarcoma. The EWSR1::FLI1 fusion, arising from an 11;22 chromosomal translocation first characterized in 1993, produces a chimeric transcription factor that requires the DNA-binding domain encoded by FLI1 for oncogenic transformation. This is one of the cleaner examples in sarcoma where a single molecular event defines disease identity.
The limits of single-marker testing: The WHO guidelines currently designate only a small number of mutational tests as critical to sarcoma diagnosis, with histological interpretation remaining the primary methodology. The key problem is that unique, single driver events characterize only a fraction of sarcoma subtypes. Modern surveys of driver distributions show that the majority of mutations exist at less than 5% frequency across the sarcoma landscape, and that potentially hundreds of unknown drivers underlie rarer subtypes. This long tail of rare molecular events is precisely what standard clinical panels miss and what pan-sarcoma genomic approaches aim to capture.
Panel sequencing results: In one pan-sarcoma study deploying a targeted panel of 465 genes plus additional informed targets, many diagnostic refinements were reported. Notably, 137 liposarcomas that had been diagnosed as "not otherwise specified" could be re-diagnosed as either well-differentiated or dedifferentiated liposarcoma based on the molecular data. More broadly, around 10% of diagnoses in that cohort could be redefined using this targeted panel alone, a finding that illustrates both the diagnostic gaps in routine histological assessment and the incremental value of even focused molecular testing.
As DNA sequencing costs dropped dramatically through the 2010s, researchers broadened the sarcoma genomics focus from specific mutations and pathways to whole-exome sequencing (WES) and whole genome sequencing (WGS), alongside transcriptional and methylation analyses. Pan-sarcoma studies aim to derive a new classification system by contrasting the many sarcoma types and subtypes within a single statistical framework. The hope is that patterns invisible to targeted panels, such as structural variants, copy number alterations, mutational signatures, and chromatin organization, can be extracted from these richer datasets and eventually incorporated into diagnostic guidelines.
Genomics England and the 100,000 Genomes Project: Genomics England Ltd, a UK Government-funded company, has sequenced more than 1,600 sarcomas as part of the 100,000 Genomes Project. The primary report, published in 2024, found that 34% of 1,617 sarcomas carried actionable mutations derived from whole genome data. This finding directly led to the commissioning of diagnostic WGS for sarcoma patients receiving treatment in NHS England centres, marking a formal transition toward multi-omics standard-of-care. A companion Cell paper published in 2025 on osteosarcoma from this programme documented ongoing chromothripsis as a driver of genome complexity and clonal evolution in that disease.
Immunogenomics: In 2024, a large multi-omics study incorporating immune profiling across 1,340 sarcomas of diverse histological subtypes revealed five cross-sarcoma immunotypes. Critically, immune depletion was found to detrimentally affect overall survival across these immunotypes, establishing immune contexture as a prognostically relevant dimension of sarcoma biology. This type of cross-sarcoma finding, integrating genomic, transcriptomic, and immunological data, exemplifies both the analytical power and the complexity of extracting actionable classifications from large multi-omics datasets. Simple clustering methods tend to fail in this setting; machine learning is increasingly required to identify biologically coherent groupings.
Since the 2010s, machine learning algorithms have been deployed across multiple coding environments and computer platforms for cancer classification tasks. Machine learning offers a distinct advantage over standard statistical analyses: complex, abstracted patterns can be derived from data that would otherwise be invisible to conventional models. For sarcoma classification specifically, the random forest method has emerged as particularly well-suited. Random forests operate through an ensemble of decision trees, each learning which features (such as the presence or absence of a specific mutation, methylation mark, or gene expression value) are most informative for distinguishing classes. The ensemble structure reduces overfitting and produces robust predictions even on noisy, high-dimensional genomic inputs.
Why methylation arrays suit this problem: Methylation arrays measure the DNA methylation state at thousands of CpG sites across the genome. Methylation patterns are a known hallmark of many sarcoma types, tend to be stable across tumor cells, and are often more consistent classifiers than mutation profiles in histologically ambiguous cases. Importantly, methylation arrays are cost-effective and bioinformatically lightweight compared to whole genome sequencing, making them deployable in many healthcare settings that cannot support a full WGS pipeline. This combination of biological signal richness and practical accessibility is what made methylation profiling the substrate for the DKFZ Sarcoma Classifier.
Precedent in CNS tumour classification: The methodological foundation of the DKFZ Sarcoma Classifier was adapted from an already published methylation array-based classifier for central nervous system tumours, published in Nature in 2018. CNS tumour classification faces comparable challenges to sarcoma, including histological ambiguity, rare subtypes, and the need to integrate molecular data with morphological assessment. The success of methylation classification in CNS pathology, which has since entered clinical practice in many neuropathology centres, provided both a proof of concept and a methodological template for the sarcoma application.
The DKFZ Sarcoma Classifier (German Cancer Research Center classifier) was developed in 2021 by groups at Heidelberg University Hospital and published in Nature Communications. The classifier uses DNA methylation profiling as its input, processed through a random forest algorithm trained to discriminate sarcoma types based on genome-wide methylation patterns. The training set comprised 1,077 sarcomas distributed across 65 types and subtypes, representing one of the largest curated sarcoma methylation datasets assembled at that time. The reported error accuracy of the classifier on its training set was 0.65%, an impressive figure reflecting strong internal performance.
Real-world validation in 428 cases: Upon real-world validation on an independent set of 428 cases, the diagnostic match between the classifier output and the institutional pathology diagnosis was 61%. In 29 of the discrepant cases, institutional diagnoses were subsequently revised following investigation of the mismatch, improving the effective match rate somewhat. This type of finding, where the classifier's output prompts a successful diagnostic revision, illustrates one of the genuine clinical use cases for AI in sarcoma pathology: not replacing the pathologist but flagging cases where standard histological interpretation may have been insufficient.
Commercial availability: The classifier has been further developed since its original publication and is now available through its commercial owner, Epignostix, at https://app.epignostix.com/#/classifiers. The existence of a commercially deployed tool based on this methodology marks an important milestone: it is one of the few AI-driven sarcoma diagnostic tools that has moved from academic research into a form accessible to clinical laboratories. This transition also raises questions about reproducibility, transparency, and the ongoing update cycle as new sarcoma subtypes are defined and added to WHO classification.
The commentary centres substantially on a validation study published in the Journal of Pathology: Clinical Research in 2021 by Lyskjaer et al, which assessed the accuracy of the DKFZ Classifier on an independent UK patient dataset. This is an important type of study because internal validation on a training cohort, regardless of how rigorous, does not reliably predict performance when a model is applied to data from different institutions, patient populations, or laboratory workflows. The complete UK validation dataset comprised 986 sarcomas, a larger cohort than the original classifier's real-world validation set of 428 cases, and critically, it included subtypes not represented in the original training set.
Overall diagnostic match: In the full 986-case dataset, 55% of cases received a diagnostic prediction from the classifier, and 450 cases (45%) had a diagnosis that matched the institutional pathology. When the analysis was restricted to only those subtypes for which the classifier was actually trained, performance improved substantially: 61% of cases received a prediction, and 88% of those predictions matched the institutional diagnosis. This distinction is methodologically important and clinically meaningful. An 88% diagnostic match rate within the classifier's trained scope represents meaningful performance, but the gap between this and the 45% overall match rate reveals the substantial coverage problem created by untrained subtypes.
The 25% coverage gap: Approximately 25% of the institutionally diagnosed subtypes in the UK dataset, including 19 distinct subtypes of soft tissue sarcoma, were not represented in the DKFZ Classifier's original training set. These cases necessarily fall outside the classifier's scope, and for patients with these subtypes, the tool provides no diagnostic output. As with the original validation, a small number of institutional diagnoses were revised following identification of discrepancies, a consistent finding across validation studies that reinforces the potential value of the classifier as a quality assurance tool even when it does not replace the primary pathological assessment.
The rare subtype performance gap: The UK validation study explicitly documented the variation in performance across individual sarcoma subtypes. For subtypes with larger and more consistent training representation, such as mesenchymal chondrosarcoma and chordoma, diagnostic match rates in the region of 90% were achievable. For rarer diseases with fewer available training data points, the results were far less reliable. Malignant peripheral nerve sheath tumours (MPNST), for example, showed a diagnostic match rate of only 24% in the validation. MPNST is a biologically aggressive subtype with significant histological overlap with other spindle cell neoplasms, and its rarity makes it difficult to accumulate training cases at any single institution.
The fundamental dataset problem: In standard machine learning practice, desirable training sets typically require thousands of labeled examples per class to achieve robust generalization. Sarcoma datasets fall very short of this threshold. Even the largest pan-sarcoma datasets assembled through national programmes contain hundreds of cases per subtype at best, and for rare subtypes, the count may be fewer than 50 or even fewer than 10. This data scarcity problem is not solvable by any single institution alone, and it creates a structural ceiling on the accuracy any single-center classifier can achieve for rare sarcoma entities.
Evolving taxonomy: A specific challenge for sarcoma AI classifiers is that the WHO classification system itself continues to evolve. An AI model trained on the 2013 WHO classification taxonomy may assign labels that are no longer considered valid under the 2020 edition. New subtypes are added, old boundaries are redrawn, and the molecular criteria for some entities change entirely. Maintaining classifier accuracy in this context requires continuous retraining and validation, which is resource-intensive and requires ongoing curatorial effort to maintain labeled training data aligned with current clinical standards.
Interpretability in pathological practice: Unlike in some medical imaging tasks where visual saliency maps can partially explain AI decisions, methylation-based classifiers present interpretability challenges. The random forest model assigns classification probabilities based on patterns across thousands of CpG sites, and it is not straightforward to translate this into the kind of biological rationale that pathologists use to communicate diagnostic reasoning or justify clinical decisions. Explainability tools such as SHAP (SHapley Additive exPlanations) values can identify which methylation features most contributed to a given prediction, but this does not provide the mechanistic narrative that is standard in pathology reporting.
Federated datasets as a structural solution: The most direct solution to the rare-subtype data scarcity problem is to compile or federate the larger datasets that exist across multiple institutions and national programmes into a single training resource of sufficient size. Federated learning approaches allow model training across multiple institutions without requiring raw data sharing, with each site computing local model updates and sharing only gradients or model weights for central aggregation. For sarcoma, where no single institution sees sufficient cases of rare subtypes, federated learning has the potential to unlock training volumes that would otherwise remain inaccessible. Several international sarcoma networks are positioned to contribute to such an effort.
Task-specific classifiers for confused diagnoses: An alternative strategic direction proposed in this commentary is to move away from all-encompassing pan-sarcoma classifiers and instead train specific algorithms to differentiate the sarcoma diagnoses that are most frequently confused in clinical practice. This approach aligns machine learning tasks more closely with the actual diagnostic problems encountered by pathologists. Pairs or small groups of morphologically overlapping entities, such as MPNST versus synovial sarcoma, or high-grade undifferentiated pleomorphic sarcoma versus dedifferentiated liposarcoma, could be addressed with targeted classifiers optimized for those specific discriminations, potentially achieving higher accuracy than a single broad-spectrum model.
Multi-modal integration: The current generation of sarcoma AI tools focuses predominantly on a single data type, most commonly methylation arrays or histopathology images. Future classifiers will need to integrate data across modalities: combining methylation profiles with transcriptomic data, targeted mutation panels, copy number alteration landscapes, and histopathological image features. Multi-modal fusion architectures, including cross-attention transformer models that can learn relationships between different data types, have demonstrated superior performance in other cancer settings. For sarcoma, where no single molecular test is sufficient for all subtypes, the combination of orthogonal data sources offers the most plausible path to comprehensive classification coverage.
Clinical integration and the regulatory pathway: The next decade will determine whether AI-driven classifiers can transition from research tools into routine clinical diagnostics, particularly for rare sarcoma subtypes where histology alone is insufficient. This transition requires not only technical accuracy but also prospective validation studies, regulatory approval pathways, integration into pathology laboratory information systems, and acceptance by the clinical and pathological communities. The DKFZ Classifier's availability through the Epignostix commercial platform represents an early step along this pathway, and it provides a template both for what is achievable and for the challenges that remain in scaling AI-driven sarcoma diagnosis.