Rapid Classification of Sarcomas Using Methylation Fingerprint: A Pilot Study

Cancers 2023 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
The Diagnostic Problem: Why Sarcoma Classification Is Notoriously Difficult

Sarcoma is a malignancy of connective tissue, divided into two broad categories: bone sarcomas and soft-tissue sarcomas (STS). Despite this seemingly simple split, the disease encompasses well over 70 histological subtypes, each with distinct molecular characteristics, clinical behavior, and treatment implications. Determining the correct subtype is not merely an academic exercise. A liposarcoma treated as an undifferentiated sarcoma, or a synovial sarcoma misclassified as a fibrosarcoma, will receive suboptimal therapy. Yet current pathological methods achieve a definitive classification in only 80-85% of cases, leaving 15-20% unclassified even after extensive workup at expert centers.

Why classification is so hard: Sarcoma diagnosis begins with biopsy and histological review, but morphological overlap between subtypes is substantial, and interobserver variability among pathologists is well-documented. Definitive subtyping often requires a layered molecular workup, including fluorescence in situ hybridization (FISH) for translocation detection, Sanger sequencing, next-generation sequencing (NGS) panels, and sometimes gene expression profiling. Institutions without access to these technologies, the majority of treating centers globally, frequently face diagnostic delays of weeks to months as samples are shipped to reference laboratories.

MDM2 amplification as a diagnostic target: A notable exception to the general lack of sarcoma-specific genomic markers is MDM2 amplification, located on chromosome 12q13-15. Elevated MDM2 expression is tightly linked to the de-differentiation process in liposarcomas, making MDM2 amplification both a diagnostic and prognostic marker for well-differentiated and dedifferentiated liposarcoma (WDLS/DDLS). Detection currently relies on FISH and immunohistochemistry (IHC), both of which require specialized equipment and expertise. The study specifically examines whether low-coverage whole-genome sequencing via nanopore can replicate this finding as a byproduct of copy-number analysis.

The clinical urgency is real: sarcoma prognosis is stage-dependent, and diagnostic delays translate directly into delayed treatment, increased risk of metastasis, and worse survival. A point-of-care tool that could deliver subtype information within hours of biopsy, rather than weeks, would represent a meaningful clinical advance.

TL;DR: Current pathological methods classify only 80-85% of sarcomas. Subtype determination requires expensive multi-step molecular workups including FISH, Sanger sequencing, and NGS, creating delays at non-expert centers. MDM2 amplification on chromosome 12q13-15 is the most established sarcoma genomic marker, specifically linked to WDLS/DDLS. This paper tests whether low-cost nanopore sequencing can classify sarcomas rapidly at the point of care.
Pages 2-3
DNA Methylation as a Tumor Fingerprint and the Nanopore Advantage

Every cell type in the body carries a characteristic pattern of DNA methylation, the addition of methyl groups to cytosine bases at CpG dinucleotide sites. These methylation patterns are stable, heritable across cell divisions, and fundamentally linked to gene expression profiles. Tumor cells retain much of the methylation signature of their tissue of origin while acquiring cancer-specific alterations. This makes genome-wide CpG methylation profiles a powerful molecular fingerprint for tumor classification, an insight that has been most thoroughly validated in central nervous system (CNS) tumors.

CNS tumor classification as the proof of concept: A landmark classifier developed at the German Cancer Research Center in Heidelberg demonstrated that DNA methylation profiles measured by Illumina BeadChip 450K arrays could reliably distinguish over 90 CNS tumor types and subtypes. This classifier has since been adopted into clinical practice in neuropathology and has substantially reduced diagnostic uncertainty for brain tumors. The key question this sarcoma study addresses is whether a similar methylation-based classification strategy can be transferred to sarcoma, a cancer type with far more heterogeneous genomic landscapes.

Why nanopore sequencing changes the equation: The Heidelberg CNS classifier and similar tools rely on Illumina Array methylation profiling, a technology with several practical drawbacks. Illumina Arrays require bisulfite conversion of DNA, a process that degrades a significant fraction of the input material and introduces technical artifacts. The assay requires centralized high-volume laboratory infrastructure and typically takes days to weeks from sample receipt to result. Oxford Nanopore Technologies sequencers, by contrast, detect methylation directly by measuring the characteristic disruptions in ionic current caused by 5-methylcytosine as DNA translocates through a protein pore. No bisulfite conversion is needed, protecting DNA integrity. The MinION device is roughly the size of a USB stick, operates at ambient conditions, and can produce usable sequencing data within hours of starting a run.

Nanopore platforms also simultaneously generate data relevant to copy-number analysis and structural variant detection (translocations), meaning a single low-coverage sequencing run can in principle deliver methylation classification, MDM2 amplification status, and fusion gene information in a unified workflow. This multi-layer output from a single assay is a key advantage over the current standard of care, which requires separate assays for each molecular question.

TL;DR: DNA methylation patterns serve as stable tumor fingerprints. The Heidelberg CNS classifier validated this concept using Illumina 450K arrays across 90+ tumor types and is now in clinical use for brain tumors. Oxford Nanopore sequencing detects methylation directly (no bisulfite conversion), runs on a compact MinION device, and can simultaneously generate copy-number and structural variant data from a single low-coverage whole-genome run.
Pages 3-5
Study Design, Cohort, Sequencing Protocol, and the nanoDx Pipeline

The study enrolled 23 patients diagnosed with sarcoma between 2018 and 2023 at Hadassah Medical Center, all of whom provided informed consent under ethics approval number 0346-12. Tissue was collected from surgically resected masses and immediately fresh-frozen, a critical requirement because the nanopore-based methylation pipeline is optimized for high-quality DNA from fresh tissue rather than formalin-fixed paraffin-embedded (FFPE) material. DNA was extracted using a DNeasy blood and tissue kit (Qiagen), quantified by Qubit fluorometry, and assessed for quality before library preparation.

Sequencing protocol: Between 200 and 400 ng of genomic DNA per sample was used for library preparation with barcode labeling via the Rapid Barcoding Kit (SQK-RBK004, Oxford Nanopore Technologies). Low-coverage whole-genome sequencing (lcWGS) was performed on a MinION Mk1C device running the R9.4.1 flow cell (FLO-MIN106D). Sequencing continued until the recommended threshold of 100 million base pairs (100M bps) per sample was reached, a coverage depth that works out to approximately 0.03-6x genome coverage depending on read quality and run performance. Of the 23 samples, 20 met the minimum coverage requirement to be processed by the downstream pipeline. Of these 20, 2 contained no tumor tissue and were excluded from the primary statistical analysis, leaving 18 tumor samples representing 11 pathologically identified sarcoma subtypes.

The nanoDx pipeline: FAST5 raw signal files and FASTQ sequence files were processed using the nanoDx pipeline (v5.0.1), an open-source workflow originally developed for nanopore methylation classification of brain tumors. The pipeline runs using snakemake (v5.4.0) workflow management on a high-performance computing (HPC) cluster. CpG methylation status was called genome-wide using nanopolish (v0.13.2), which assigns binary methylation values (1 for methylated, 0 for unmethylated) to each detected CpG site by analyzing the ionic current signal. Methylation frequency per site is computed as the fraction of reads classified as methylated, providing a continuous value between 0 and 1.

Adapting the pipeline for sarcoma: The original nanoDx pipeline uses the Heidelberg CNS reference cohort as its training dataset. The critical methodological modification in this study was the substitution of a custom sarcoma reference set built from Illumina Array methylation data covering 62 sarcoma methylation classes. This single pipeline adjustment, replacing the reference cohort while keeping all other parameters including Random Forest hyperparameters identical to the CNS version, forms the core innovation of the study. Copy-number profiles were generated from the same sequencing run by aligning FASTQ reads to the hg19 human reference genome using minimap2 (v2.15) and binning reads into 1000 kilobase pair (kbp) bins using QDNAseq via R/Bioconductor. MDM2 amplification was assessed from these copy-number profiles.

TL;DR: 23 patients, 20 sequenced successfully, 18 tumor samples from 11 sarcoma subtypes analyzed. lcWGS on MinION Mk1C using R9.4.1 flow cells, targeting 100M bps per sample. nanoDx pipeline (v5.0.1) with nanopolish for CpG calling, snakemake workflow on HPC. Key modification: replaced CNS reference cohort with a custom 62-class sarcoma reference set built from Illumina Array data. Copy-number profiles generated simultaneously from same run using minimap2 and QDNAseq.
Pages 5-6
Random Forest Classifier: Architecture, Confidence Scoring, and t-SNE Validation

The methylation-based classification engine at the heart of the nanoDx pipeline is a Random Forest classifier implemented in Python using the RandomForestClassifier function from scikit-learn (v1.0.2). Random Forest is an ensemble learning method that builds hundreds of decision trees, each trained on a random subset of features (in this case, CpG methylation frequency values at specific genomic loci). The final class prediction is the majority vote across all trees, and the class probability can be estimated from the fraction of trees voting for each class. For a dataset of 50,000 or more CpG sites, Random Forest is computationally tractable and provides reasonable performance even with limited training examples per class, a practical advantage in a rare cancer setting.

Confidence scoring: Raw Random Forest class probabilities are not well-calibrated as confidence measures. The pipeline addresses this by applying the CalibratedClassifierCV function from scikit-learn, which recalibrates the class probability estimates to more accurately reflect the true likelihood of correct classification. The output is reported as a "confidence score." In the CNS brain tumor classifier, a confidence score threshold of 0.15 was established as the minimum for a reliable classification call. This study applies the same 0.15 threshold to sarcoma data without recalibration for the sarcoma context, a methodological choice the authors explicitly acknowledge as a limitation. The "Random Forest voting rate" refers to the fraction of individual trees that voted for the top classification class and is reported as a complementary metric alongside the confidence score.

t-SNE unsupervised clustering as orthogonal validation: In addition to the supervised Random Forest classifier, the pipeline performs unsupervised clustering using t-SNE (t-distributed stochastic neighbor embedding) applied to the 50,000 most variable CpG sites across all samples. t-SNE projects high-dimensional methylation data into a two-dimensional scatter plot where samples with similar methylation profiles cluster together. Agreement between t-SNE cluster membership and Random Forest classification provides a form of cross-validation that does not rely on the training reference set. Importantly, t-SNE is not used as a classifier, it is explicitly a quality control and visualization tool.

A key observation is that cases with low confidence scores consistently showed disagreement between the Random Forest classification and t-SNE clustering, providing a practical internal consistency check. All samples with a confidence score below 0.08 in this cohort produced discordant classifications, which the authors suggest may define a lower threshold for valid classification in the sarcoma setting.

TL;DR: Random Forest classifier from scikit-learn (v1.0.2) with CalibratedClassifierCV confidence scoring. CNS-derived threshold of 0.15 confidence applied without sarcoma-specific recalibration. t-SNE on 50,000 most variable CpG sites used as orthogonal visual quality control. All cases with confidence score below 0.08 produced discordant classifications, suggesting a sarcoma-specific lower validity threshold may be needed.
Pages 6-8
Classification Performance: 78% Concordance, Copy-Number Findings, and a Striking Discordant Case

The Random Forest classifier achieved a 78% concordance rate (14 of 18 samples) with the pathological diagnosis. The 18 tumor samples included a mean read length of 3,966 base pairs (range 1,310 to 7,078 bp), with a mean of 27,594 CpG sites covered per sample (range 6,000 to 70,000 sites). Genome coverage ranged from approximately 0.04x to 6.4x depending on the sample. For the 14 concordantly classified cases, the median confidence score was 0.18 (range 0.08 to 0.88), comparable to the confidence scores reported in the validated CNS nanopore classifier. Half of the concordant cases exceeded the 0.15 CNS threshold for reliable classification. The mean Random Forest voting rate for the most confident concordant classifications was 11.23% (range 4.4 to 27.6%).

Discordant cases: Four samples (22%) showed discordant classification. Three of the four had very low confidence scores, ranging from 0.04 to 0.07, well below the 0.08 threshold the authors propose as potentially indicative of a non-valid classification for sarcoma. These three cases included two myxofibrosarcoma samples (SARC-12 and SARC-22) and one chondrosarcoma (SARC-13). The fourth discordant case, SARC-07, is particularly informative. The classifier assigned it to Ewing sarcoma (EWS) with a confidence score of 0.91 and a voting rate of 30%, both the highest in the entire cohort. The pathology report classified SARC-07 as a small round blue cell tumor (SRBCS) without EWS translocation on FISH. However, Ewing sarcoma and CIC-rearranged sarcomas, a closely related entity in the SRBCS family, share overlapping methylation profiles, and FISH for EWS translocation can be negative in CIC-rearranged tumors. The authors interpret this case as a possible genuine Ewing-like sarcoma misclassified by conventional pathology rather than a pipeline error.

Copy-number and MDM2 amplification: Five of the 18 samples (28%) showed MDM2 amplification detectable on the copy-number profile derived from the same nanopore sequencing run. All five were classified under the WDLS/DDLS methylation class by the Random Forest, confirming that MDM2 amplification co-segregates with the liposarcoma methylation signature. In three cases with low confidence scores (range 0.08 to 0.14) that fell below the 0.15 reliability threshold, the presence of MDM2 amplification provided an independent genomic corroboration of the methylation-based classification, demonstrating the additive value of combining copy-number information with the methylation classifier.

t-SNE concordance: Among the 14 concordantly classified tumors, 12 (86%) also clustered correctly in the t-SNE analysis, providing strong orthogonal support for the classifier's accuracy. For the 4 discordant cases, 2 also disagreed with t-SNE clustering (both had confidence scores of 0.04), while SARC-07 (confidence 0.91) did agree with t-SNE, further supporting the hypothesis that the EWS classification is genuinely correct.

TL;DR: 78% concordance (14/18) with pathology. Median confidence score for concordant cases: 0.18 (range 0.08-0.88). Mean read length 3,966 bp, mean CpG sites covered 27,594. Five samples (28%) showed MDM2 amplification, all co-classified as WDLS/DDLS. All cases below confidence 0.08 were discordant. t-SNE agreed with classifier in 12/14 concordant cases (86%). Intriguing discordant case SARC-07 had confidence 0.91 and EWS t-SNE clustering, suggesting possible pathology misclassification.
Pages 8-10
Interpreting the Results: What 78% Concordance Means in a Rare Cancer Pilot

A 78% concordance rate in a pilot study of 18 heterogeneous sarcoma tumor samples, spanning 11 distinct pathological subtypes, achieved using an unmodified brain tumor classifier with only its reference dataset swapped, is a meaningful proof of concept rather than a benchmark for clinical deployment. The authors are careful to frame their results in this context, repeatedly noting that no Random Forest hyperparameters were adjusted for sarcoma and that the CNS-derived confidence threshold of 0.15 was applied without recalibration. Both of these factors mean that the classifier is likely underperforming relative to what would be achievable with sarcoma-specific optimization.

The significance of the single pipeline change: The sole methodological adaptation was replacing the Heidelberg CNS reference cohort with the authors' in-house sarcoma reference set covering 62 methylation classes. The fact that this single substitution, with no other changes, yielded 78% concordance across 11 sarcoma subtypes in a small prospective cohort illustrates the generalizability of the methylation fingerprinting approach and the power of the nanoDx pipeline architecture. It implies that extending this method to other cancer types with suitable reference sets is technically feasible without requiring fundamentally new algorithmic development.

The SARC-07 case as a microcosm of the diagnostic utility: The most compelling result in the paper is arguably the discordant SARC-07 case. A small round blue cell tumor without detectable EWS FISH translocation would classically be left without a definitive molecular diagnosis. The methylation classifier assigned it to Ewing sarcoma with exceptional confidence (0.91), and this assignment was corroborated by t-SNE clustering. CIC-rearranged sarcomas and other Ewing-like tumors share the small round blue cell morphology and can be FISH-negative for the canonical EWSR1-FLI1 translocation. The possibility that the methylation classifier correctly identified a diagnosis that the standard FISH workup missed is scientifically plausible and clinically consequential, though it requires prospective validation.

The authors also discuss the practical implications of the confidence score behavior. Low confidence scores correlate reliably with incorrect or unreliable classification, providing a built-in quality filter: a classifier output with confidence below 0.08 can be treated as indeterminate, analogous to the "not otherwise classifiable" category in conventional pathology. This property is important for clinical safety because it gives the classifier a mechanism to decline to make a call rather than imposing a wrong answer.

TL;DR: 78% concordance across 11 subtypes was achieved with only a reference dataset swap, no hyperparameter changes. SARC-07 (confidence 0.91, Ewing sarcoma call) may represent a genuine pathology misclassification of a CIC-rearranged sarcoma. Confidence scores below 0.08 reliably flag invalid classifications, functioning as a built-in indeterminate filter. Sarcoma-specific hyperparameter optimization is expected to improve performance further.
Pages 10-11
Key Limitations: Sample Size, Reference Set, Fresh Tissue Dependency, and Unoptimized Parameters

Small cohort and narrow subtype coverage: The most significant limitation is the cohort size. Eighteen tumor samples spanning 11 sarcoma subtypes means that most subtypes are represented by one or two cases, making it statistically impossible to estimate per-subtype accuracy or generalize findings across the full spectrum of sarcoma pathology. Sarcoma encompasses 70+ recognized subtypes, and the 11 included in this study represent only a fraction of the diagnostic challenge. Subtypes prone to misclassification, such as epithelioid sarcoma, alveolar soft part sarcoma, and desmoplastic small round cell tumor, were not represented in the cohort.

Unoptimized hyperparameters: The Random Forest classifier used in this study was built with hyperparameters validated for CNS tumor classification, including the minimum number of CpG sites required for model training and the confidence score threshold. No sarcoma-specific calibration was performed. The authors explicitly note that rescaling these parameters to the sarcoma context would likely improve classification performance and more accurately reflect the true reliability of confidence scores in this disease setting. The practical implication is that the 78% concordance rate is a lower bound on what the approach can achieve once properly calibrated.

Fresh tissue requirement: The nanopore methylation pipeline performs best with high-molecular-weight DNA from freshly frozen tissue. FFPE tissue, which is the standard archival format for most pathology material and the most accessible sample type in clinical practice, suffers from formalin-induced DNA crosslinking and fragmentation that degrades methylation signal quality. This fresh tissue dependency restricts prospective clinical use to cases where fresh surgical material is available, excluding retrospective studies and settings where immediate freezing protocols are not in place.

Reference set limitations: The sarcoma reference set used to train the classifier comprised 62 methylation classes but was derived from Illumina Array data rather than nanopore data. Cross-platform differences in CpG site coverage, methylation calling algorithms, and technical noise introduce potential discordance between the reference set and the test samples. The reference set will need to be expanded, ideally with nanopore-derived reference profiles, as more samples are collected. Additionally, with only 62 reference classes, several sarcoma subtypes are likely absent from the reference, meaning the classifier cannot even nominally assign those subtypes.

TL;DR: 18 tumor samples across 11 subtypes is too small for per-subtype accuracy estimates. Hyperparameters not recalibrated for sarcoma (CNS-derived thresholds used). Fresh frozen tissue required, excluding FFPE samples. Reference set of 62 classes built from Illumina Array data, not nanopore, creating cross-platform noise. All limitations suggest current results are a floor, not a ceiling, on what the approach can achieve.
Pages 11-12
Path to Clinical Integration: Multimodal Classifiers, Broader Validation, and Point-of-Care Potential

Expanding the reference set and sarcoma-specific calibration: The most immediate priority identified by the authors is growing the sarcoma reference set beyond 62 classes and recalibrating the Random Forest hyperparameters for sarcoma-specific data. As more institutions adopt nanopore sequencing for research purposes, multi-institutional sarcoma methylation databases will enable construction of larger, more representative reference cohorts. A nanopore-native reference set, rather than one derived from Illumina Array data, would eliminate cross-platform noise and likely improve classification confidence for borderline cases.

Integrating additional molecular layers: The authors envision a future classifier that combines methylation profiles with transcriptomic data (RNA sequencing), proteomic information, pathological morphological features, and full copy-number alteration landscapes including not only MDM2 but also other recurrent sarcoma genomic events such as CDK4 amplification, CDKN2A deletion, and ATRX loss. Point mutations detectable from the same nanopore run could be incorporated as additional classification features. This multi-omics integration is analogous to the trajectory seen in CNS tumor classification, where the WHO 2021 classification now formally requires molecular information alongside histology for definitive diagnosis, and a sarcoma equivalent may be achievable.

Liquid biopsy extension: The authors specifically note that the methylation and copy-number information generated by nanopore sequencing could be extended to cell-free DNA from plasma, enabling non-invasive monitoring of sarcoma treatment response and minimal residual disease through liquid biopsy. Fragmentomics (the analysis of cell-free DNA fragment length patterns) and methylation analysis from plasma DNA have both been explored in other cancer types with promising results. Translating this to sarcoma would require significant validation but offers the prospect of monitoring patients between imaging timepoints without requiring repeat tissue biopsy.

Point-of-care deployment and intraoperative application: The MinION's compact form factor and rapid data generation raise the possibility of intraoperative classification during tumor resection. A surgeon could obtain classification information within the duration of a surgical procedure, potentially guiding real-time decisions about resection margins or the need for intraoperative lymph node sampling. While this is a longer-term aspiration requiring substantial workflow and regulatory development, the technical barriers are lower for nanopore than for any alternative sequencing platform currently available. The authors conclude that the demonstrated feasibility across a restricted range of 11 sarcoma subtypes provides sufficient justification for a larger multicenter prospective validation study.

TL;DR: Next steps include expanding the reference set beyond 62 classes, sarcoma-specific hyperparameter calibration, and adding transcriptomic, proteomic, and mutation data to the classifier. Liquid biopsy application for treatment monitoring via plasma cell-free DNA is a stated future direction. Intraoperative real-time classification during surgery is a longer-term aspiration enabled by MinION's compact form factor. A multicenter prospective validation study is the immediate priority.