Osteosarcoma is the most common primary malignant bone tumour in children and young adults. The current standard of care begins with neoadjuvant combination chemotherapy using high-dose methotrexate, doxorubicin, and cisplatin (the MAP regimen), followed by surgical resection and adjuvant chemotherapy. Despite significant advances in multi-agent chemotherapy since the 1980s, 5-year event-free survival rates for localised disease have remained largely unchanged at 60 to 70%, dropping to only 30% for metastatic disease.
Histological response as the gold standard: Assessing whether a tumour responded to pre-surgical chemotherapy currently requires examining the resected specimen under a microscope after surgery. A threshold of 90% or more tumour necrosis on histopathology is widely accepted as indicating a favourable response and is used for risk stratification in clinical trials. However, this benchmark is only available after surgery, limiting its utility for real-time treatment decisions. Whether histological response is truly prognostic remains debated, with some evidence suggesting its predictive value is affected by induction chemotherapy intensity.
Why MRI volume changes matter: MRI-based tumour volume changes have been proposed as a non-invasive, presurgical biomarker of chemotherapy response. Early research demonstrated a correlation between MRI-derived three-dimensional volume measurements and histopathologic necrosis. However, the methods used in routine clinical practice, specifically three-diameter measurement and the ellipsoid formula (V = pabc/6), suffer from high inter-observer variability of up to 25% and intra-observer variability of 6 to 14%, substantially limiting their reliability as response biomarkers.
Manual MRI tumour segmentation, the more accurate alternative to the ellipsoid formula, is time-consuming and requires specialised expertise, making it impractical for routine use. This study addresses that gap by developing and validating two automated convolutional neural networks (CNNs): one for tumour volume segmentation and one for therapy response prediction, validated across two independent academic medical centres.
This retrospective, multicentre study was conducted at two academic institutions: Centre A (Stanford University) and Centre B (Children's Wisconsin). IRB approval was obtained at Centre A (IRB 48854) with patient consent waived given the retrospective design. The study enrolled patients under 26 years of age with biopsy-proven high-grade osteosarcoma, requiring availability of MRI scans at both baseline and after MAP chemotherapy, performed within 4 months of each other.
Centre A (training and internal validation): 112 patients were initially identified, with 21 excluded for missing histopathological necrosis reports (n = 3), tumour surgery before chemotherapy or more than 4 weeks after the second MRI (n = 4), incomplete MRI scans (n = 12), or severe imaging artefacts (n = 2). The final Centre A cohort comprised 91 patients (39 females, 52 males, mean age 15 plus or minus 5 years, range 6 to 26 years). Of these, 81 patients (162 annotated MRI scans) formed the training set and 10 patients formed the internal validation set. The femur was the most common primary tumour site (51%), followed by the tibia (20%). Exactly 54% of Centre A patients achieved the favourable histologic response threshold of 90% or more necrosis.
Centre B (external validation): 10 patients (2 females, 8 males, mean age 13 plus or minus 0 years) provided 20 MRI scans for external validation. MRI data were collected between January 2018 and July 2024. The tibia was the dominant tumour site at Centre B (70%), and only 30% of patients achieved the 90% necrosis threshold, compared with 54% at Centre A. This difference in responder prevalence between centres reflects the limited sample size at Centre B and the challenge of external generalisability.
Chemotherapy timing and imaging protocol: All patients received neoadjuvant MAP chemotherapy over a mean of 9 plus or minus 3 weeks. Post-chemotherapy MRI scans were performed at a mean of 10 plus or minus 4 weeks after baseline, with local control surgery at 15 plus or minus 10 weeks. MRI protocols varied across centres and scanner models, including 3 T GE, Siemens, and Philips platforms, with imaging sequences including Gadolinium-enhanced T1-weighted, fat-saturated T2-FSE, diffusion-weighted imaging (b-values 50 and 600 to 800 s/mm2), and fat-saturated post-contrast T1 sequences.
The volume estimation algorithm processes Gadolinium-enhanced T1-weighted MR images to generate three-dimensional tumour volume measurements in cubic millimetres. The pipeline begins with a pre-conditioning step that standardises image properties across scanners and protocols, a critical step for handling the heterogeneous acquisition parameters across the two centres.
Segmentation CNN architecture: The CNN was trained on radiologist-delineated contours generated by a reader with seven years of subspecialty experience in oncologic imaging. The model architecture contains 12 convolutional layers with ReLU activations and 31.1 million parameters (118.69 MB). Training used a combined cross-entropy and Dice loss function with the Adam optimiser. The input was 128 x 128 x 128 voxel patches, with 0.5/0.5/0.5 subsampling to balance geometric fidelity against dataset size constraints. Manual segmentation was performed in the axial plane using 3D Slicer software, with Gadolinium-enhanced images superimposed with plain T1 and T2 sequences to improve tumour delineation.
Outcome prediction model: The therapy response prediction pipeline incorporated two independent instances of the volume estimation CNN, one applied to the baseline MRI and one to the post-chemotherapy MRI. The change in tumour volume between these two time points served as the primary imaging biomarker. This delta volume was then fed into a logistic regression outcome prediction module trained with radiologist-labelled imaging features (baseline volume, post-induction volume, and tumour volume change) alongside histopathological necrosis labels as the ground truth. Patients were classified as responders (90% or more necrosis) or non-responders (less than 90% necrosis).
Inter-observer and intra-observer reliability: Two board-certified radiologists each performed independent segmentations, both with two-year subspecialty focus in oncologic imaging. Inter-observer variability between reader 1 (307 plus or minus 595 mm3) and reader 2 (328 plus or minus 761 mm3) showed moderate agreement with an intraclass correlation coefficient (ICC) of 0.7 (p = 0.001). Intra-observer agreement for reader 1, tested on 10 cases after a 2-month interval, was high (ICC = 0.9, p = 0.001), confirming that a single expert's repeated measurements are highly consistent even when the tumour's CNN-based and ellipsoid-formula volumes are compared.
The volume estimation CNN achieved Dice coefficients of 0.9 at Centre A (internal validation) and 0.8 at Centre B (external validation), with mean Hausdorff distances of 15 plus or minus 19 mm and 14 plus or minus 8 mm, respectively. The Dice coefficient measures spatial overlap at the voxel level between the CNN-predicted segmentation and the radiologist-defined ground truth, with 1.0 indicating perfect agreement. Values of 0.8 to 0.9 are generally considered strong performance for tumour segmentation tasks, comparable to published deep learning models for meningiomas and non-small cell lung cancer (NSCLC).
Internal validation (Centre A): AI-derived and human-derived tumour volumes showed strong Spearman correlation (r = 0.98, p less than 0.001). Bland-Altman analysis revealed a mean bias of negative 6 mm3 with 95% limits of agreement from negative 112 to positive 99 mm3, indicating minimal systematic deviation. By comparison, volumes calculated using the traditional ellipsoid formula versus manual segmentation showed a mean bias of 115 mm3 with 95% limits from negative 351 to positive 580 mm3, a substantially larger spread that quantifies the inaccuracy of the ellipsoid approach.
External validation (Centre B): The CNN maintained strong performance on the independent external cohort, with AI versus human correlation of r = 0.95 (p less than 0.001). Bland-Altman analysis showed a mean bias of negative 0.2 mm3 with 95% limits from negative 13 to positive 12 mm3, an exceptionally narrow agreement interval. The ellipsoid formula versus human segmentation comparison at Centre B yielded a mean bias of 7 mm3 with 95% limits from negative 21 to positive 35 mm3. The CNN's performance at Centre B was achieved despite differences in scanner hardware, field strength (1.5 T and 3 T), and imaging protocols relative to the training data, demonstrating meaningful cross-scanner generalisation.
Tumours without histological response (less than 90% necrosis) typically exhibited significant volume growth post-therapy on the segmented MRI scans, while those achieving 90% or more necrosis showed size reduction or stability. This visual pattern was captured quantitatively by the delta volume metric used in the outcome prediction module.
Due to the limited sample size for external validation, patients from Centre A's internal validation set and Centre B's external validation cohort were combined for response stratification analysis, yielding 20 cases total. Histopathology confirmed 90% or more necrosis in 12 of 20 osteosarcomas and less than 90% necrosis in 8 of 20. Post-induction MRI scans demonstrated stable or decreased tumour volume in 12 of 20 cases and increased tumour volume in the remaining 8, with this imaging-based grouping aligning with the histopathology distribution.
Classification results: The CNN-based outcome prediction model correctly classified 16 of 20 cases, achieving 80% diagnostic accuracy, 90% sensitivity (correctly identifying responders), and 70% specificity (correctly identifying non-responders). In comparative terms, the human reader (Reader A) applying the same volumetric classification logic correctly classified 15 of 20 cases, giving the CNN a slight advantage over experienced radiologist assessment. This comparison indicates the CNN matches or slightly exceeds expert human performance on this task, a notable result given the small validation cohort.
Clinical interpretation of sensitivity and specificity: The 90% sensitivity means the model correctly identified 9 out of 10 true responders, minimising the risk of withholding favourable prognosis information from patients who actually responded well. The 70% specificity means 3 out of 10 non-responders were incorrectly labelled as responders. In clinical practice, this asymmetry may be acceptable if the goal is to identify patients who might benefit from treatment escalation, since false negatives (missed non-responders) carry greater consequence than false positives in this context.
This performance compares favourably to prior work. Teo et al. used Fuzzy C-Means clustering and weighted majority voting on a cohort of only 10 patients for a similar prediction task. Bouhamama et al. developed an MRI-based radiomics model for preoperative risk stratification achieving AUC 0.95, though that approach required manual radiomic feature engineering rather than the end-to-end CNN pipeline used in this study.
The discussion contextualises these results within a broader landscape of AI applications in oncologic imaging. For osteosarcoma specifically, prior published AI work has largely focused on alternative imaging modalities. Hasei et al. developed an X-ray AI model for osteosarcoma diagnosis achieving sensitivity 95.5%, specificity 96.2%, and AUC 0.99, though plain radiographs provide less volumetric information than MRI. Huang et al. introduced a CT-based segmentation method achieving Dice of 87.8%. Ouyang et al. proposed UATransNet, an MRI-based U-Net incorporating self-aware attention mechanisms, also achieving a Dice of 0.9, directly comparable to this study's internal validation performance.
Benchmark comparisons from other tumour types: Deep learning segmentation models for meningiomas have achieved accuracy comparable to human experts (Dice approximately 0.9). In NSCLC, 3D-Slicer-based AI contouring has demonstrated stronger correlation with pathology and reduced inter-observer variability. Kawaguchi et al. showed AI-derived solid tumour volumes in stage I lung adenocarcinoma improved survival prediction (AUC 0.8). These cross-tumour comparisons validate that the Dice and correlation metrics reported in this osteosarcoma study are consistent with the state of the art across oncologic imaging applications.
The paediatric AI gap: AI applications in paediatric bone tumours are notably undercharacterised relative to adult oncology, largely due to the rarity of these tumours and the inherent challenges of assembling large paediatric imaging datasets. Published paediatric AI work has focused primarily on brain tumours, with a small number of studies addressing lymphoma, neuroblastoma, Wilms tumour, and neuroendocrine tumours. Response prediction or outcomes modelling in paediatric sarcoma specifically remains largely unexplored territory, making this study one of the first to address it directly with validated CNN models.
The comparison to human reader performance (15/20 vs. 16/20 correct classifications) is particularly meaningful for establishing clinical credibility. AI models that match or exceed expert human performance on objective metrics provide a more compelling case for integration into clinical workflows than those evaluated only against algorithm benchmarks.
Retrospective design and selection bias: The study's retrospective nature introduces potential selection bias. Patients with incomplete MRI data, missing histopathology reports, or scans acquired outside the protocol window were excluded, which may mean the final cohort over-represents cases with clean imaging and straightforward tumour boundaries. Real-world prospective deployment would encounter noisier and more incomplete data than the curated retrospective cohort.
Small external validation cohort: The external validation set comprised only 10 patients from Centre B. While statistically significant agreement was still demonstrated, the statistical power to detect meaningful differences in model performance between subgroups (by tumour site, scanner type, or histologic response threshold) is limited. The authors appropriately note that the generalisability of findings should be interpreted with caution, and the small size necessitated pooling Centre A internal validation and Centre B external validation patients for the response prediction analysis.
Training data limited to two scanner platforms: CNNs trained on MRI scans from two centres may not generalise reliably to imaging protocols at other institutions, particularly those using different pulse sequences, contrast agents, or field strengths outside the training distribution. The inclusion of 1.5 T and 3 T scanners from multiple manufacturers mitigates this somewhat, but single-centre AI models have a consistent track record of performance drops when tested externally.
Tumour volume as the sole response biomarker: The model evaluates chemotherapy response exclusively through changes in tumour volume. In paediatric oncology, favourable responses can manifest as architectural changes, including internal necrosis, cystic degeneration, and changes in signal intensity, without necessarily producing major volume reductions. A tumour that undergoes significant internal necrosis while maintaining its outer dimensions would be misclassified as a non-responder by a volume-only model. Future iterations should incorporate functional MRI parameters such as diffusion-weighted imaging (apparent diffusion coefficient), texture features, and internal signal heterogeneity.
Incorporating functional MRI parameters: The next generation of osteosarcoma response prediction models should extend beyond anatomical volume changes to include functional imaging biomarkers. Diffusion-weighted imaging (DWI) quantifies tumour cellularity through the apparent diffusion coefficient (ADC), which has been correlated with histologic necrosis in prior osteosarcoma studies. Textural radiomic features capturing signal heterogeneity within the tumour volume could identify cases where internal composition changes without gross volume change, a response pattern the current model would miss. Dynamic contrast-enhanced MRI perfusion parameters are also candidates for inclusion.
Integration with clinical and molecular variables: Future models should combine imaging biomarkers with clinical variables such as initial tumour volume, patient age, tumour site, and lactate dehydrogenase levels, as well as molecular markers where available. The MAP regimen intensity has been reported to affect the prognostic value of histological necrosis, suggesting that treatment-related variables should be incorporated as model inputs rather than held constant across training cases.
Prospective multicentre validation: The study authors explicitly call for prospective, multicentre datasets with larger and more diverse patient populations to improve model robustness and clinical applicability. A prospective design would also enable the use of event-free survival and overall survival as outcome endpoints rather than histopathologic necrosis alone. Given that histological response remains a debated surrogate endpoint in osteosarcoma, validating the CNN's predictions against long-term survival data would strengthen the clinical case for deployment.
Clinical workflow integration: Automated volumetric analysis could be embedded into radiology reporting workflows to generate standardised pre- and post-chemotherapy tumour volume reports without requiring manual segmentation. This would make 3D volumetric response assessment practical in centres that currently rely on the ellipsoid formula due to resource constraints. The reduced inter-observer variability of the CNN (Bland-Altman limits of agreement of negative 112 to positive 99 mm3) compared to the ellipsoid formula (negative 351 to positive 580 mm3) already makes a quantitative case for replacing the ellipsoid formula in routine practice. Combined with the NCI grant funding (R01 CA269231) supporting this research, these tools are positioned for continued development toward clinical deployment.