Feasibility of Multi-Parametric MRI Combined with Machine Learning in the Assessment of Necrosis of Osteosarcoma After Neoadjuvant Chemotherapy: A Preliminary Study

BMC Cancer 2020 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Why Assessing Chemotherapy Response in Osteosarcoma Is So Difficult

Osteosarcoma is the most common primary malignant bone tumor in children and adolescents, and it ranks among the leading causes of cancer-related death and disability in these age groups worldwide. The standard treatment paradigm relies heavily on neoadjuvant chemotherapy (NACT) before surgery, a strategy that has substantially improved the 5-year survival rate and enabled limb-salvage procedures in many patients who would otherwise require amputation. However, osteosarcoma is biologically heterogeneous, and patients respond very differently to the same chemotherapy regimen.

The core problem: The most reliable indicator of NACT effectiveness is the tumor necrosis rate measured from the surgically resected specimen, typically assessed according to the Huvos grading system. A tumor necrosis rate of 90% or greater (Huvos grade III-IV) is associated with significantly better overall survival and guides decisions about postoperative chemotherapy intensity. The critical limitation is that this assessment can only be performed after surgery. There is no established, noninvasive, preoperative method to accurately quantify tumor necrosis while the patient is still receiving chemotherapy.

Current imaging limitations: Magnetic resonance imaging (MRI) is the primary modality for local staging and postoperative surveillance of bone tumors. Conventional MRI parameters including tumor volume, marrow edema, and enhancement patterns have been evaluated for chemotherapy response assessment, but none have achieved the accuracy needed to reliably replace histological grading. Diffusion-weighted imaging (DWI) with apparent diffusion coefficient (ADC) mapping offers a quantitative window into tissue microstructure, and several studies have shown that ADC values rise as tumor cells die after effective chemotherapy. However, ADC alone still has limitations, particularly when it comes to cartilaginous components of the tumor.

This 2020 preliminary study from the First Affiliated Hospital of Sun Yat-Sen University asks whether combining multiple MRI parameters with machine learning can overcome the limitations of single-parameter imaging assessment. The hypothesis is that multi-parametric MRI (mpMRI), feeding several quantitative imaging features simultaneously into a machine learning classifier, can more accurately distinguish viable from nonviable tumor tissue across all histological subtypes, including the diagnostically challenging cartilaginous components.

TL;DR: Osteosarcoma NACT response is currently evaluated only from post-surgical specimens using the Huvos grade, where a necrosis rate of 90%+ predicts better survival. ADC from DWI-MRI can track tumor cell death noninvasively, but struggles with cartilaginous tumor components. This study tests whether combining ADC with T2-weighted and contrast-enhanced T1-weighted MRI signals, fed into a random forest classifier, improves necrosis discrimination preoperatively.
Pages 2-3
Prospective Cohort, MRI Protocol, and Tissue Sampling Design

This was a prospective study enrolling 12 patients (7 males, 5 females; mean age 14.6 plus or minus 4.8 years, range 6-25 years) with primary osteosarcoma of the limbs, recruited consecutively at a single academic center between August 2011 and March 2012. While the sample size is small, the prospective design with rigorous histological correlation represents a significant methodological strength over retrospective imaging-only studies. All patients underwent four cycles of NACT with high-dose methotrexate, pirarubicin, and ifosfamide, with or without cisplatinum, followed by limb-salvage surgery approximately 3 weeks after completing chemotherapy.

Tumor locations and histological subtypes: Eight patients had osteosarcoma of the distal femur and four had proximal tibia involvement. The histological breakdown was osteoblastic (n=7, 58.3%), chondroblastic (n=4, 33.3%), and fibroblastic (n=1, 8.4%). The inclusion of chondroblastic cases is particularly relevant because cartilaginous components pose a specific challenge for ADC-based assessment, as viable cartilaginous tumor cells paradoxically produce high ADC values similar to necrotic tissue, creating diagnostic ambiguity that standard DWI cannot resolve.

MRI acquisition: All imaging was performed on a Siemens Magnetom Trio 3.0 Tesla whole-body MRI scanner using an extremity coil. The protocol included T1-weighted imaging (T1WI) with TR=659 ms and TE=11 ms, T2-weighted imaging (T2WI) with fat suppression at TR=4660 ms and TE=96 ms, and axial DWI using a single-shot spin-echo echo-planar imaging sequence at b-values of 0 and 800 s/mm2. Subtract-enhanced T1WI (ST1WI), which isolates contrast enhancement by subtracting pre-contrast from post-contrast T1 images, was also acquired with gadolinium contrast. Slice thickness was 5 mm with a 1 mm interslice gap, with field of view centered on the largest tumor cross-section.

Tissue-to-image coregistration: MRI was performed within 3 days before surgery, and resected gross specimens were fixed in 10% buffered formaldehyde and sectioned into 5 mm axial slices matching the MRI imaging planes. Radiologist-pathologist coregistration selected 6-10 well-matched specimen sections per patient. Rectangular tissue samples ranging from 10x15 mm to 15x20 mm were drawn on these sections, yielding 127 total tissue samples from the 12 patients, of which 102 were classified as histologically homogeneous and included in the analysis.

TL;DR: 12 patients (mean age 14.6 years, 7 male), prospectively recruited, with distal femur (n=8) and proximal tibia (n=4) osteosarcoma. All received 4-cycle NACT before limb-salvage surgery. 3.0T MRI acquired ADC (DWI b=0, 800 s/mm2), T2WI with fat suppression, and subtracted contrast-enhanced T1WI. 127 tissue samples obtained via rigorous MRI-to-gross-specimen coregistration; 102 homogeneous areas used for analysis.
Pages 3-4
The Five-Category Tissue Framework and Why Cartilaginous Viability Is the Hard Problem

The study's pathological classification scheme defines five distinct tissue categories observed in post-NACT osteosarcoma specimens. Viable tumor tissue is divided into two types: non-cartilaginous viable areas, containing living sarcomatous cells, tumor osteoid, and tumor bone, and cartilaginous viable areas, containing viable chondrosarcomatous cells within a cartilaginous matrix. Nonviable (necrotic) tissue is divided into three categories: non-cartilaginous tumor necrosis (sarcomatous cell necrosis), post-necrotic collagen (fibrotic replacement of dead tumor), and tumor necrotic cystic or hemorrhagic areas including secondary aneurysmal bone cysts (ABC).

Viability threshold: Areas with less than 10% tumor cell necrosis were classified as viable, while those with 90% or greater necrosis were classified as nonviable. The 25 samples with intermediate necrosis (10-90%, the partial necrosis zone) were deliberately excluded from analysis to avoid ambiguous ground truth labels that would compromise classifier training and evaluation. This is a methodologically sound decision, though it means the classifier was not tested on the clinically important borderline cases.

Sample distribution: Of the 102 homogeneous tissue samples, 38 (37.3%) were non-cartilaginous viable tumor, 25 (24.5%) were non-cartilaginous tumor necrosis, 14 (13.7%) were cartilaginous viable tumor, 14 (13.7%) were tumor necrotic cystic or hemorrhagic areas including secondary ABC, and 11 (10.8%) were post-necrotic collagen. The relatively small number of cartilaginous viable samples (n=14) represents both the clinical reality and a key limitation, as it constrains the statistical power of the cartilaginous subclassification task.

Why cartilaginous tissue is a problem for ADC: Normal cartilage contains abundant water bound within proteoglycan-collagen matrix, which restricts water diffusion but in a different pattern from hypercellular solid tumor. Viable chondrosarcomatous cells in a cartilaginous matrix produce high ADC values, which overlap substantially with the high ADC values seen in necrotic cystic or collagenized areas. This makes it impossible for ADC alone to distinguish living cartilaginous tumor from dead tumor matrix, an overlap that does not occur with non-cartilaginous osteosarcoma components where hypercellularity reliably restricts diffusion.

TL;DR: Five tissue categories were defined: non-cartilaginous viable (n=38, 37.3%), cartilaginous viable (n=14, 13.7%), non-cartilaginous necrosis (n=25, 24.5%), cystic/hemorrhagic necrosis and secondary ABC (n=14, 13.7%), and post-necrotic collagen (n=11, 10.8%). 25 partially necrotic samples were excluded. ADC cannot distinguish cartilaginous viable from necrotic tissue because both yield high diffusion values, motivating the multi-parametric approach.
Pages 4-5
Feature Extraction, Normalization, and Random Forest Classification

For each of the 102 tissue samples, circular or oval regions of interest (ROIs) ranging from 50 to 250 mm2 were placed on T2WI, subtract-enhanced T1WI (ST1WI), and ADC maps, coregistered to the histological sampling areas by two experienced musculoskeletal radiologists working jointly. The average MRI signal within each ROI was measured for three parameters: ADC, T2WI signal intensity, and ST1WI signal intensity. To control for inter-patient variability in absolute MRI signal levels, each parameter was normalized by dividing by the corresponding signal from normal adjacent muscle, yielding standardized ratios: rADC (standardized ADC), rT2WI (standardized T2-weighted signal), and rST1WI (standardized subtract-enhanced T1-weighted signal).

Three binary classification tasks: Rather than attempting a five-class problem with limited data, the authors defined three binary classification tasks of increasing difficulty. Task 1 distinguished non-cartilaginous tumor survival from tumor nonviable (the most tractable problem, since ADC already works well here). Task 2 distinguished all tumor survival (both cartilaginous and non-cartilaginous) from all tumor nonviable (the clinically most relevant question for computing overall necrosis rate). Task 3 distinguished cartilaginous tumor survival specifically from tumor nonviable (the hardest problem, where ADC fails).

Random forest algorithm: All three classifiers used the random forest (RF) algorithm, an ensemble method that builds a large number of decision trees during training and outputs the majority vote across trees as the class prediction. Random forests handle small datasets well, are inherently resistant to overfitting through the bootstrap aggregation and random feature selection that differentiates each tree, and produce probabilistic outputs that can be used to compute receiver operating characteristic (ROC) curves. Models were trained and evaluated using leave-one-out cross-validation (LOOCV) given the small sample size, meaning each sample was held out in turn as a test point while the remaining samples trained the model.

Comparative evaluation: For each classification task, the team built two RF models side by side: one using rADC as the only input feature, and one using the full multi-parametric combination of rADC, rT2WI, and rST1WI. Performance was compared using AUC, sensitivity, specificity, and accuracy, with statistical significance of AUC differences tested using the DeLong method. This two-model comparison structure directly quantifies the incremental value of adding T2WI and contrast enhancement information to ADC alone.

TL;DR: ROIs (50-250 mm2) placed on ADC, T2WI, and ST1WI maps coregistered to histological samples. All signals normalized to adjacent normal muscle (rADC, rT2WI, rST1WI). Three binary RF classifiers trained for: non-cartilaginous viable vs. nonviable, all viable vs. nonviable, and cartilaginous viable vs. nonviable. Leave-one-out cross-validation used given n=102. Two models compared per task: rADC alone vs. rADC plus rT2WI plus rST1WI.
Pages 5-7
Classification Performance: Where mpMRI Makes a Difference

The results demonstrate a consistent pattern: multi-parametric MRI improves classification accuracy over ADC alone in all three tasks, with the improvement reaching statistical significance in the two tasks involving cartilaginous tissue. For the most tractable task, distinguishing non-cartilaginous tumor survival from tumor nonviable, the rADC-only model achieved a sensitivity of 88%, specificity of 89%, accuracy of 89%, and AUC of 0.93. Adding rT2WI and rST1WI pushed sensitivity to 97%, specificity to 96%, accuracy to 96%, and AUC to 0.97. The AUC difference of 0.04 was not statistically significant (P=0.0933), indicating that ADC alone already performs very well for non-cartilaginous discrimination.

All-tissue viable vs. nonviable: For the clinically central task of distinguishing all tumor survival from all tumor nonviable, the rADC model yielded sensitivity 82%, specificity 69%, accuracy 75%, and AUC 0.83. The mpMRI model improved significantly to sensitivity 94%, specificity 78%, accuracy 85%, and AUC 0.90, with the AUC difference reaching statistical significance (P=0.0473). This is the most clinically meaningful result, because this task mirrors the overall necrosis rate calculation used in the Huvos grading system. The specificity gap (69% vs. 78%) is particularly notable, as false positives (calling viable tumor nonviable) would lead to underestimating treatment response.

Cartilaginous viable vs. nonviable: The hardest classification task confirmed the expected weakness of ADC alone. The rADC-only model for distinguishing cartilaginous tumor survival from tumor nonviable achieved sensitivity 68%, specificity 57%, accuracy 66%, and AUC of only 0.61, which is barely above chance. The mpMRI model with all three parameters yielded sensitivity 66%, specificity 92%, accuracy 71%, and AUC of 0.81, a statistically significant AUC improvement of 0.20 (P=0.0153). The dramatic specificity gain (57% to 92%) reflects that T2WI and ST1WI signals help the classifier correctly identify nonviable tissue as nonviable, reducing false positives where dead cartilaginous matrix is misclassified as living tumor.

MRI parameter values by tissue type: The standardized parameter values reveal the biological basis for discrimination. Non-cartilaginous viable tumor had the lowest rADC (0.94 plus or minus 0.16), consistent with hypercellularity restricting diffusion. Cartilaginous viable tumor had a much higher rADC (1.53 plus or minus 0.13), overlapping with post-necrotic collagen (1.82 plus or minus 0.13) and cystic or hemorrhagic necrosis (1.76 plus or minus 0.21). The rT2WI signal separated cartilaginous viable tumor (5.86 plus or minus 1.54) from non-cartilaginous viable tumor (3.81 plus or minus 1.84), and rST1WI (contrast enhancement) provided additional discriminatory information that ADC alone cannot supply.

TL;DR: Non-cartilaginous viable vs. nonviable: rADC AUC=0.93, mpMRI AUC=0.97 (P=0.0933, not significant). All viable vs. nonviable: rADC AUC=0.83 vs. mpMRI AUC=0.90 (P=0.0473, significant). Cartilaginous viable vs. nonviable: rADC AUC=0.61 (near chance) vs. mpMRI AUC=0.81 (P=0.0153, significant), with specificity improving from 57% to 92%. mpMRI provides the largest benefit where ADC is least useful: cartilaginous tumor components.
Pages 7-8
Why Each MRI Parameter Contributes Different Biological Information

The discussion section provides biological grounding for the multi-parametric approach. ADC reflects tissue cellularity and membrane integrity through water molecule diffusion. In non-cartilaginous viable tumor, high cellularity and intact cell membranes create crowded extracellular spaces that restrict water diffusion, resulting in low ADC values. After effective chemotherapy, cell death disrupts membrane integrity and reduces cellularity, allowing water molecules to move more freely and raising ADC. This mechanism explains why ADC works well for non-cartilaginous components, where the viable-to-nonviable transition produces a clear, directional ADC shift.

Why cartilage undermines ADC-only analysis: Cartilaginous tissue, whether viable chondrosarcomatous cells or acellular cartilaginous matrix, maintains high ADC values because the sparse proteoglycan-collagen matrix of cartilage allows relatively free water diffusion even in living tissue. This means the ADC of viable cartilaginous tumor is already elevated at baseline, overlapping with the elevated ADC of post-necrotic collagen and cystic spaces. There is essentially no chemotherapy-induced ADC shift to detect. T2WI signal intensity, by contrast, reflects water content and structural organization of the extracellular matrix, and can differentiate the hyperintense cartilaginous matrix from other tissue types. Contrast enhancement (ST1WI) reflects tissue vascularity and endothelial permeability, properties that differ between viable, vascularized tumor tissue and avascular, post-necrotic stroma.

Multi-parametric imaging in other cancers: The authors draw on a broader literature showing that multi-parametric imaging with different functional MRI parameters provides richer information about cancer biology than any single parameter. In prostate cancer, the PI-RADS classification system already mandates combining T2WI, DWI, and dynamic contrast enhancement (DCE) for lesion assessment. Studies in rectal cancer and breast cancer have similarly demonstrated that mpMRI outperforms single-parameter approaches for treatment response evaluation. The osteosarcoma context is analogous, with the additional complexity that osteosarcoma is histologically heterogeneous within a single tumor, containing areas of bone, cartilage, fibrous tissue, and necrosis in varying proportions.

The combination of rADC, rT2WI, and rST1WI captures complementary aspects of tumor biology: diffusion restriction from cellularity, T2 relaxation reflecting tissue water organization, and perfusion-related enhancement from vascularity. Together they allow the random forest classifier to build decision boundaries that neither imaging feature could establish alone, particularly in the cartilaginous subspace where ADC is fundamentally uninformative.

TL;DR: ADC tracks cellularity changes in non-cartilaginous tumor but fails for cartilaginous components because viable cartilage already has high ADC at baseline. T2WI differentiates cartilaginous matrix from other tissue types through water-matrix interactions. ST1WI captures vascularity differences between viable and post-necrotic tissue. Together these three parameters span independent biological dimensions, enabling the random forest to overcome ADC's fundamental limitation in cartilaginous osteosarcoma.
Pages 8-9
Small Cohort, Single Center, and the Gap Between Tissue Classification and Clinical Necrosis Rate

Small sample size and preliminary scope: The most prominent limitation is the cohort of 12 patients producing 102 analyzable tissue samples. The study explicitly labels itself a "preliminary study," and the authors acknowledge that the findings require validation in a larger, multicenter prospective cohort before any clinical application can be considered. The small sample size constrains the statistical power of the cross-validation estimates, and LOOCV on 102 samples, while appropriate given the data volume, may still produce optimistic accuracy estimates compared to true external validation. The 95% confidence intervals on sensitivity, specificity, and AUC are correspondingly wide.

Tissue-level vs. patient-level assessment: A critical conceptual gap exists between the study's output and clinical decision-making. The classifiers operate at the level of individual tissue samples, distinguishing whether a given 10-15 mm rectangular sample is predominantly viable or nonviable. The Huvos grading system, by contrast, requires a whole-tumor necrosis rate expressed as a percentage of the entire resected specimen. The study does not demonstrate whether aggregating tissue-sample-level classifier outputs across the whole tumor can produce an accurate overall necrosis rate estimate. This translation from sample-level classification to patient-level response assessment is the key missing step.

Single-center, single-scanner design: All imaging was performed on a single 3.0T Siemens scanner with a standardized protocol, and all specimens were processed at a single institution. The rADC, rT2WI, and rST1WI features are computed relative to adjacent muscle, which partly controls for between-scanner differences in absolute signal levels, but scanner manufacturer, field strength, pulse sequence parameters, and contrast agent dosing all affect quantitative MRI values in ways that normalization does not fully resolve. Multi-center validation would be required to assess whether the classifier thresholds learned at one institution generalize to others.

Exclusion of heterogeneous samples: The 25 tissue samples with intermediate necrosis rates (10-90%) were excluded from analysis to ensure clean ground truth labels. In clinical practice, however, tumors contain exactly this type of partially necrotic heterogeneous tissue. A classifier trained only on clearly viable or clearly dead tissue may not perform reliably on the mixed-response tissue that pathologists most need guidance about. Additionally, the exclusion reduces the effective dataset size and may introduce selection bias.

TL;DR: Key limitations are the small cohort (12 patients, 102 samples), single-center single-scanner design, and the unresolved gap between tissue-sample-level classification and patient-level Huvos necrosis rate estimation. 25 partially necrotic samples were excluded, which limits applicability to heterogeneous tumors. External multicenter validation and a method for aggregating sample-level outputs into a whole-tumor necrosis rate are the two most critical next steps.
Pages 9-10
From Proof of Concept to Clinical Decision Support: What Needs to Happen Next

Patient-level necrosis rate estimation: The most urgent next step is extending the tissue-level classifier to whole-tumor response assessment. This would require segmenting the entire tumor volume on each MRI series, applying the classifier across all voxels or regions within that volume, and computing the proportion classified as nonviable as a proxy for the Huvos necrosis rate. The threshold for defining a "good responder" (90% necrosis by Huvos) would need to be calibrated on a larger dataset with matched postoperative pathology. Radiomics pipelines capable of whole-tumor feature extraction and voxel-wise classification have been developed for other sarcoma subtypes and could be adapted here.

Longitudinal response monitoring: The current study acquires MRI only at a single time point, immediately before surgery. A more clinically powerful application would be serial mpMRI at baseline and during chemotherapy (for example, after 2 cycles of NACT), enabling early identification of non-responders while there is still time to modify the treatment plan. Machine learning models trained on changes in rADC, rT2WI, and rST1WI between time points could quantify the trajectory of necrosis induction, potentially providing actionable guidance before the course of therapy is complete.

Advanced machine learning architectures: This study used classical supervised random forest classification on hand-crafted ROI-level features. Future work could apply convolutional neural networks (CNNs) directly to co-registered MRI volumes, learning spatial features from the entire tumor without requiring manual ROI placement or explicit feature engineering. Deep learning models trained on multi-channel MRI inputs (stacking T1WI, T2WI, DWI, and CE-MRI as separate input channels) have been applied successfully to treatment response prediction in other cancers, including rectal and cervical cancer. Transfer learning from pre-trained radiology foundation models could help overcome the small dataset constraint.

Integration with clinical and genomic data: MRI-based response assessment in isolation captures only one dimension of treatment response. The significant heterogeneity of osteosarcoma biology suggests that integrating MRI-derived features with pretreatment tumor biopsy genomics (including copy number profiles and gene expression of chemotherapy resistance pathways), serum biomarkers such as alkaline phosphatase and lactate dehydrogenase, and patient-level pharmacokinetic data could produce more accurate composite response predictors. Multimodal machine learning frameworks capable of fusing imaging, genomic, and clinical tabular data are an active area of development in computational oncology and represent the long-term direction for precision sarcoma care.

TL;DR: Key next steps include: extending tissue-level classification to whole-tumor Huvos necrosis rate estimation via volumetric segmentation; serial mpMRI at baseline and during chemotherapy for early non-responder identification; applying CNN-based deep learning to multi-channel MRI volumes rather than hand-crafted ROI features; and multimodal integration of imaging with biopsy genomics and serum biomarkers for comprehensive response prediction in osteosarcoma.