Chondrosarcoma (CHS) is a malignant bone tumor defined by the production of cartilaginous matrix without typical bone mineralization. It accounts for approximately 30% of all malignant bone tumors, ranking second in frequency after osteosarcoma, with an annual incidence of roughly five cases per million people. Unlike many solid tumors, chondrosarcoma is largely resistant to conventional cytotoxic chemotherapy and radiation, making complete surgical resection the primary curative intervention. This chemoresistance is a defining clinical challenge: once a tumor is unresectable or has metastasized, treatment options are severely limited.
Clinical heterogeneity: Chondrosarcoma encompasses several distinct histological subtypes with markedly different biological behaviors. Conventional (NOS) chondrosarcoma is the most common, but dedifferentiated chondrosarcoma, mesenchymal chondrosarcoma, myxoid chondrosarcoma, clear cell chondrosarcoma, and malignant chondroblastoma each carry different prognoses and may require distinct management strategies. Tumor grade further stratifies risk: Grade I (well differentiated) tumors can often be managed with curettage, while Grade III-IV (poorly to undifferentiated) tumors carry high rates of local recurrence and distant metastasis with substantially worse overall survival.
The prognostic gap: Traditional staging and prognostic tools, most notably the TNM system and basic nomograms built on logistic regression or Cox proportional hazard models, summarize group-level risk but do not flexibly capture the interaction among multiple prognostic factors for individual patients. Existing nomograms produce a single point estimate and cannot dynamically show how prognosis changes at different follow-up horizons (1-year vs. 5-year vs. 10-year survival). This motivates the development of prediction tools that can generate personalized survival probabilities at multiple clinically meaningful timepoints.
This study by Yang, Yang, Bao, and colleagues from Yunnan, China, uses the C5.0 decision tree algorithm within a recursive partitioning analysis (RPA) framework to construct separate models predicting survival status at 12, 24, 60, and 120 months. The analysis is based on the SEER database and includes 2,409 to 3,894 patients depending on the time point, making it one of the larger machine learning studies in chondrosarcoma. The final models were translated into both a desktop software application and a publicly accessible web-based Shiny platform to facilitate real-world clinical use.
The data source for this study is the Surveillance, Epidemiology, and End Results (SEER) database, administered by the National Cancer Institute (NCI). SEER is a nationwide cancer registry covering approximately 34% of the U.S. population, and contains patient demographics, tumor characteristics, treatment information, and longitudinal survival data. The authors extracted records for patients with a definitive diagnosis of chondrosarcoma as identified by ICD-O-3 codes, diagnosed between 2000 and 2018. This 19-year window was chosen to ensure sufficient sample size while avoiding older, potentially less complete coding practices.
Inclusion and exclusion: Patients were included if they had a confirmed chondrosarcoma diagnosis in the 2000-2018 period with complete baseline characteristics, histological subtype, tumor grade, treatment modalities, and documented survival status at the relevant time point. Patients diagnosed before 2000 were excluded, as were those with ambiguous or missing survival duration and those with censored records during the follow-up interval. Censored records were excluded because the C5.0 algorithm requires a definitive binary outcome (alive vs. dead) at each time point rather than a survival time estimate. This exclusion led to four separate cohorts: 3,894 patients with 12-month status, 3,674 with 24-month status, 3,147 with 60-month status, and 2,409 with 120-month status.
Prognostic variables: The feature set included age at diagnosis, sex (male or female), race (white, black, or other), year of diagnosis (before 2010 vs. 2010 or later), histological subtype (7 categories including chondrosarcoma-NOS, dedifferentiated, mesenchymal, myxoid, juxtacortical, clear cell, and malignant chondroblastoma), tumor grade (Grades I through IV), primary tumor site, tumor size as a continuous variable, surgical procedure (performed vs. not), radiotherapy (yes vs. no), chemotherapy (yes vs. no), lymph node metastasis, and visceral metastasis to brain, liver, or lung. This multi-variable feature set reflects the clinical reality that chondrosarcoma prognosis is determined by the interaction of tumor biology, anatomic extent, and therapeutic response rather than any single factor alone.
C5.0 algorithm and model construction: The C5.0 algorithm is an updated version of the classic C4.5 decision tree method, operating in SPSS Modeler. It recursively partitions the patient population by selecting the feature at each node that provides the greatest information gain ratio, splitting patients into progressively more homogeneous subgroups. The resulting tree structure visualizes decision pathways where each branch represents a specific combination of prognostic features and each terminal leaf node carries a predicted survival probability. Each cohort was randomly split into a 70% training set and a 30% test set. Model performance was evaluated using accuracy rate (ACC) and the area under the receiver operating characteristic (ROC) curve (AUC) in both training and test datasets. The study was reported according to the TRIPOD statement for multivariable prediction model research.
The baseline characteristics reveal a patient population with mean age of 52.8 years (SD 18.5 years) in the 12-month cohort, consistent with the known epidemiology of chondrosarcoma as a disease affecting adults across a wide age range with peak incidence in the fifth and sixth decades. The sex distribution showed a modest male predominance, with males comprising 54.3% of the 12-month cohort and a similar proportion across all four cohorts. The racial composition was predominantly white (86.5%) across all cohorts, reflecting the known higher incidence of chondrosarcoma in white populations and the demographic coverage of the SEER registry.
Survival rates by time horizon: The observed survival probabilities paint a clear picture of chondrosarcoma's long-term trajectory. At 12 months, 87.3% of patients were alive, reflecting the relatively slow-growing nature of many chondrosarcomas in the short term. At 24 months, survival dropped to 80.0%, and by 60 months only 67.1% remained alive. The 10-year (120-month) survival rate of 48.4% underscores that while many patients survive short-term, approximately half ultimately die from their disease over a decade of follow-up. These survival rates are consistent with population-based studies and highlight the importance of long-horizon prediction tools rather than relying solely on short-term prognostication.
Tumor characteristics: The mean tumor size across the 12-month cohort was 82.8 mm (SD 55.9 mm), reflecting the typically large size at diagnosis for a disease that often presents insidiously. Surgery was performed in 85.3% of patients, reinforcing that resection remains the cornerstone of treatment. In contrast, only 14.5% received radiotherapy and 8.1% received chemotherapy, consistent with the well-established radioresistance and chemoresistance of conventional chondrosarcoma. The relatively low rates of adjuvant treatment also reflect clinical practice where surgery alone is considered adequate for lower-grade tumors.
The shift in prognostic factor rankings across the four time horizons is one of the study's most clinically informative findings. Certain factors dominate early survival prediction while different factors become more influential at 5 or 10 years, reflecting the changing biology of disease progression over time. This temporal dynamism is a key advantage of the multi-horizon modeling approach.
The C5.0 algorithm identified seven variables as the most influential prognostic indicators across the four time horizons: tumor histology, surgical procedure, age, visceral metastasis to the brain/liver/lung, chemotherapy, tumor grade, and sex. These seven factors emerged as the foundation nodes in the decision tree models and consistently ranked as the highest-importance predictors. The selection of these factors by an algorithm that evaluates information gain ratio provides a data-driven validation of clinically intuitive risk factors, but also adds quantitative precision to the relative weighting of each factor.
Tumor histology and grade: These two biological factors ranked among the top predictors at all time horizons. Histological subtype determines the tumor's inherent aggressiveness: dedifferentiated chondrosarcoma, which contains a high-grade sarcomatous component within a lower-grade cartilaginous tumor, carries the worst prognosis of all subtypes with 5-year survival rates below 20% in most series. Conventional NOS chondrosarcoma spans grades I through III, with Grade I tumors behaving near-benignly and Grade IV (undifferentiated) tumors behaving like high-grade sarcomas. The decision trees capture these biologically meaningful distinctions as branching logic.
Surgery and visceral metastasis: Whether a patient underwent surgery was consistently among the most important predictors for short-term survival (12 and 24 months), reflecting that resectability directly determines whether curative intent can be achieved. Visceral metastasis to brain, liver, or lung was a dominant predictor at all time horizons, as distant organ involvement indicates advanced disease with very limited treatment options. The combination of no surgery and visceral metastasis places patients in the worst prognostic subgroup, with modeled survival probabilities dropping dramatically at each time point.
Temporal dynamics of predictor importance: A notable finding is that the ranking of predictor importance shifts across time horizons. For 12-month survival, tumor histology, surgery, and age dominate. For 60 and 120-month survival, chemotherapy and tumor grade become relatively more prominent, potentially because long-term outcomes are more shaped by the inherent aggressiveness of the tumor biology rather than the immediate surgical decision. This shift challenges the conventional practice of assessing prognostic factors only in aggregate across the full follow-up period and suggests that clinical risk stratification should be tailored to the relevant time horizon.
Model performance was evaluated using two complementary metrics: accuracy rate (ACC), which measures overall classification correctness, and the area under the ROC curve (AUC), which measures discriminatory power independent of classification threshold. Performance was assessed separately in the training and test datasets to enable a meaningful assessment of generalizability. The use of a separate held-out test set (30% of each cohort) is a more rigorous validation approach than internal cross-validation alone, as it provides a single unbiased performance estimate on data the model never encountered during training.
Training set performance: Accuracy rates in the training set ranged from 82.40% (60-month model) to 89.08% (12-month model). The corresponding AUC values were 0.806 (60-month), 0.820 (120-month), 0.834 (24-month), and 0.894 (12-month). The relatively higher performance of the 12-month model is expected because short-term mortality events are more readily predictable from clinical variables and exhibit stronger signals compared to long-term survival over a decade.
Test set performance: In the independent test dataset, accuracy ranged from 82.16% to 88.74%, and AUC values ranged from 0.808 to 0.882. Critically, the test set performance was closely aligned with training performance, with differences of less than 1% in accuracy and less than 0.02 in AUC at each time point. This minimal degradation from training to test performance suggests the models are not substantially overfitting to the training data. Confusion scatter plots, which display predicted vs. actual survival status, showed tight clustering around the diagonal, confirming high concordance between model predictions and observed outcomes.
To contextualize these results: an AUC of 0.882 for the 12-month model represents substantially better discrimination than chance (0.5) and is competitive with other machine learning models published for bone sarcoma prognosis, which typically report AUC values in the range of 0.75-0.90. The 60-month model's slightly lower AUC of 0.808 reflects the inherent difficulty of 5-year predictions in a disease where treatment practices, competing risks, and new therapeutic options can shift outcomes over a long follow-up window.
A central contribution of this study is its translation of the decision tree models into an accessible clinical tool. Rather than presenting static decision tree graphics (which require clinicians to manually trace through branching logic), the authors developed both a desktop software application and a web-based interactive platform using the Shiny package within R version 4.3.3. These tools remove the burden of manual tree navigation and allow any clinician with internet access to generate personalized survival predictions for individual patients.
Application structure: The Shiny application is organized into five distinct tab pages. The "Introduction" tab provides an overview of the application's purpose and methodology. The "Datasets" tab allows users to review and perform automated visual data analysis on the underlying modeling datasets. The "Model" tab presents the constructed decision trees and the process used to derive them. The "Performance" tab displays the confusion scatter plots and ROC curves for each time horizon. The "Survival Prediction" tab is the core clinical interface, where users enter individual patient characteristics and receive calculated survival probability estimates at 12, 24, 60, and 120 months simultaneously.
Worked clinical example: The paper provides a specific worked example to illustrate how clinicians would use the tool. A 60-year-old male patient with dedifferentiated (Grade IV, undifferentiated) chondrosarcoma and confirmed visceral metastasis to the brain, liver, or lung, who has undergone surgical intervention and chemotherapy, is projected to have survival probabilities of 62.86% at 12 months, 56.13% at 24 months, 28.81% at 60 months, and 26.14% at 120 months. This type of multi-horizon estimate enables clinicians to communicate realistic prognosis trajectories to patients, distinguish short-term from long-term risk, and tailor treatment intensity and follow-up intensity accordingly.
The developers designed the platform with several intended clinical use cases: personalized treatment protocol selection based on risk stratification, interactive patient education about disease trajectory, clinical decision support at multidisciplinary tumor board discussions, and integration with institutional clinical workflow. User feedback mechanisms and data security measures were built into the design. The authors acknowledge that formal usability testing and prospective clinical integration studies are needed before the tool can be considered ready for routine deployment.
Retrospective design and SEER limitations: This study is entirely retrospective, drawing from a registry that was not designed for the purpose of prognostic model development. Retrospective analyses are subject to selection bias, ascertainment bias, and the absence of important clinical variables not collected in registry format. The SEER database does not capture detailed surgical margin status, specific chemotherapy regimens (only whether chemotherapy was given), molecular or genetic markers (such as IDH1/2 mutations, which are present in a substantial fraction of conventional chondrosarcomas and are now being targeted therapeutically), functional performance status, or comorbidity burden. These omitted variables are clinically meaningful and their absence may limit the precision of the model for individual patients.
Exclusion of censored patients: A methodological limitation inherent to the binary classification framework is the exclusion of patients with censored survival data. Patients who were lost to follow-up or who had not yet reached the relevant time point were excluded from each cohort. This introduces potential survivorship bias, as censored patients may differ systematically from those with complete follow-up. Traditional survival analysis methods such as Kaplan-Meier analysis and Cox regression handle censoring explicitly through partial likelihood estimation, which is a statistical advantage they hold over the binary decision tree approach used here.
Missing data and coding variability: SEER coding practices evolved over the study period (2000-2018), and some variables (particularly histological subtype and tumor grade) may have been coded inconsistently across registries and over time. Missing data, which is a recognized challenge in large registry studies, was not described in detail, and the approach to handling incomplete records was not fully transparent. The proportion of cases with missing key variables could meaningfully affect the representativeness of the modeling cohorts.
External generalizability: The SEER database predominantly captures U.S. patients and overrepresents white patients (86.5% in this cohort). Chondrosarcoma management practices, including surgical approaches, adjuvant treatment decisions, and surveillance strategies, may differ internationally. Models trained on SEER data may not accurately predict outcomes in populations treated under different healthcare systems or with different tumor biology distributions. The authors call for prospective clinical studies to validate the findings in independent cohorts, including non-U.S. populations and institutions with different practice patterns.
Prospective validation: The authors explicitly recommend prospective clinical studies as the necessary next step to validate the decision tree models. Retrospective registry-based training, however large, does not capture the full clinical complexity of chondrosarcoma management. Prospective studies would allow collection of variables not available in SEER, including surgical margin status, molecular mutation profiles, response to emerging targeted agents (such as ivosidenib for IDH-mutant chondrosarcoma), functional outcome measures, and quality of life endpoints. Prospective data would also allow proper handling of censored observations through survival analysis rather than binary outcome classification.
Molecular and genomic integration: The identification of recurrent IDH1 and IDH2 mutations in conventional chondrosarcoma, as well as other driver events in dedifferentiated and mesenchymal subtypes, has opened pathways to targeted therapy. The SEER database cannot capture these molecular features, but future models incorporating genomic data could substantially improve prognostic precision. Circulating tumor DNA (ctDNA) measurement with IDH1/2 and GNAS mutations is emerging as a biomarker for risk stratification in central chondrosarcoma, and integrating liquid biopsy data into machine learning models is a natural extension of the approach used in this paper. MRI and CT radiomics-based signatures for chondrosarcoma (which have shown AUC values of 0.75-0.90 in recent studies) could also be combined with clinical predictors in a multimodal prognostic framework.
Model refinement and comparison: The study compared the C5.0 decision tree approach favorably against TNM staging and existing nomograms in the chondrosarcoma literature, but did not directly compare it against other machine learning approaches such as random forests, gradient boosted machines, or neural networks on the same dataset. Future work could systematically benchmark C5.0 against these alternatives to determine whether tree-based interpretability comes at a cost in predictive performance relative to ensemble methods. The SORG (Skeletal Oncology Research Group) algorithm, which was previously developed and externally validated for chondrosarcoma 5-year survival prediction, provides a specific benchmark the authors could use in future comparative analyses.
The broader vision articulated by the authors is a user-centric platform integrating personalized treatment protocol generation, educational modules, decision support, and data feedback into existing clinical workflows. The path from a validated research model to a clinically deployed tool requires regulatory clearance, integration with electronic health record systems, usability testing with clinicians, liability framework development, and ongoing performance monitoring as practice patterns evolve. These implementation challenges are common to all clinical AI tools but are especially relevant for rare cancers where the evidence base remains thin compared to common solid tumors.