Unlocking the Power of Benchmarking: Real-World-Time Data Analysis for Enhanced Sarcoma Patient Outcomes

Cancers 2023 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Why Sarcoma Care Needs Systematic Benchmarking

Sarcomas are a heterogeneous group of malignant tumors arising from mesenchymal tissues, encompassing over 100 histological subtypes affecting bone, soft tissue, and visceral organs. Because of their rarity and biologic diversity, sarcomas demand intensive multidisciplinary management. Surgery remains the cornerstone of curative intent, but outcomes depend heavily on factors including surgical margin quality, preoperative planning, access to specialist centers, and the rigor of tumor board deliberations. Despite the existence of international guidelines, the actual practice of sarcoma care varies considerably between institutions and countries, and no systematic benchmark for comparing these practices existed prior to this work.

The quality gap: An international consensus jury, as cited in this paper's framework, identified the need for reproducible, objective, and universal outcome parameters to establish meaningful benchmarks in sarcoma surgery. Current approaches typically rely on clinical trial data collected under controlled conditions, which may not reflect the full complexity of real-world sarcoma management. Key quality dimensions such as multidisciplinary team (MDT) meeting structure, biopsy practices, radiation planning, chemotherapy indications, and patient-reported outcomes are rarely captured and compared across centers in a harmonized way.

The value-based healthcare imperative: This study is situated within the broader framework of value-based healthcare (VBHC), a model introduced by Porter et al. that defines value as the ratio of health outcomes to cost. Under traditional fee-for-service models, providers are incentivized to maximize volume rather than optimize outcome. VBHC realigns these incentives by measuring outcomes that matter to patients across the full care cycle and linking reimbursement and performance assessment to those outcomes. For sarcoma, achieving VBHC requires an interoperable data infrastructure capable of capturing clinical, patient-reported, and economic dimensions longitudinally.

The authors argue that artificial intelligence and machine-learning approaches will be critical to extracting actionable evidence from the large, heterogeneous datasets that sarcoma care generates. However, before ML models can be applied meaningfully, a structured, standardized data framework is needed. This paper presents that framework and the first proof-of-concept comparison using it.

TL;DR: Sarcoma care lacks systematic benchmarking despite international guidelines, leading to unmeasured variation in quality and cost. This paper introduces a digital platform for real-world-time (RWT) benchmarking grounded in value-based healthcare (VBHC) principles, and positions it as the prerequisite infrastructure for AI and ML applications in sarcoma care quality improvement.
Pages 2-4
Study Design: Prospective Multi-Center Data Collection Across Two MDT/SBs

The study enrolled 983 patients diagnosed with sarcoma who were consecutively presented at two independent multidisciplinary team/sarcoma board (MDT/SB) centers, designated MDT/SB-A and MDT/SB-B, over a prospective 15-month collection window. Both centers qualify as internationally representative sarcoma referral units, each processing more than 100 newly diagnosed sarcoma patients per year and operating within a tertiary university hospital network. A key methodological strength is that data were collected prospectively and consecutively, meaning no retrospective selection bias was introduced: every patient discussed at each MDT/SB was included in the dataset.

MDT/SB requirements: At both centers, formal MDT/SB attendance requires participation from all 8 relevant disciplines, including orthopedic or surgical oncology, medical oncology, radiation oncology, pathology, radiology, plastic and reconstructive surgery, rehabilitation medicine, and palliative care. Pathology reference review and comprehensive imaging review are mandatory prerequisites for case presentation. Newly diagnosed patients, patients completing each treatment step (such as after preoperative radiation and before surgery), and patients experiencing any change in treatment plan are required to be presented at the MDT/SB. This structure ensures that the data captured reflects deliberated, multidisciplinary decision-making rather than individual clinician judgments.

Statistical approach: Descriptive statistics were used for baseline characteristics, with categorical variables presented as counts and proportions and numerical variables as median with range. Fisher's exact test was applied to assess statistically significant differences in chemotherapy and biopsy rates between the two centers. The R statistical program was used for all analyses. The simplicity of the statistical toolkit is intentional: this study is designed as a proof-of-concept for the benchmarking platform rather than as a clinical hypothesis-testing trial.

The patient cohort included 452 females (46.0%), with a median age at diagnosis of 58 years (range 1 to 59 years). Metastatic disease at diagnosis was present in 44 patients (4.5%). Tumor characteristics including histological dignity and anatomic location were systematically recorded across all cases.

TL;DR: 983 sarcoma patients prospectively and consecutively enrolled across two high-volume tertiary MDT/SBs over 15 months. 46% female, median age 58 years, 4.5% metastatic at diagnosis. Both centers require all 8 disciplines and mandatory pathology and imaging review. Fisher's exact test used to compare biopsy and chemotherapy rates between centers.
Pages 4-6
Architecture and Workflow of the Sarconnector Digital Platform

The Sarconnector is an interoperable digital platform developed by the study authors (BF and PH, Zurich, Switzerland) specifically to enable real-world-time data collection, analysis, and benchmarking across sarcoma centers. It is designed around two core layers: a front end that handles data entry and real-time data visualization, and a back end built on SQL database architecture with the R statistical program for analytical processing. Data can be introduced via hospital servers, cloud servers, API data exporter tools, or interactive Shiny applications, making the platform adaptable to different institutional IT infrastructures.

Data dimensions: The Sarconnector captures four distinct data dimensions: clinician-reported outcome measures (CROMS), patient-reported outcome measures (PROMS), patient-reported experience measures (PREMS), and quality indicators (QIs). CROMS encompass clinical metrics reported by treating physicians, including surgical margin status, complication grades using the Clavien-Dindo classification and the Comprehensive Complication Index (CCI), radiation parameters such as critical tumor volume and gross tumor volume (CTV/GTV), and chemotherapy details. PROMS and PREMS capture the patient's own perspective on functional outcomes and care experience across the long-term follow-up cycle. Health economics data constitutes a fifth dimension, enabling cost analysis alongside clinical outcome assessment.

Meta-level versus ground-level data: A conceptually important distinction the authors draw is between ground-level (object-level) data and meta-level data. Ground-level data describes individual patients and their outcomes. Meta-level data, which is the Sarconnector's primary contribution, describes the structure, organization, and patterns across populations and institutions. This meta-level perspective allows the platform to identify systematic differences in care practices between MDT/SBs, rather than only describing individual case outcomes. The platform can analyze quality indicators for a single MDT/SB or integrate data across multiple centers to generate national or continental benchmarks.

Automated analysis pipeline: The Sarconnector's workflow follows a structured eight-step pipeline: (A) collection of RWTD over the full care cycle, (B) secure storage on the interoperable platform, (C) automated analysis and front-end display, (D) benchmarking against predefined sarcoma-specific quality indicators, (E) assessment of care quality from benchmark results, (F) VBHC value assessment relating outcomes to resource use, (G) iterative improvement loop, (H) predictive AI/ML modeling on the accumulated RWTD, leading ultimately to (I) composition of the sarcoma patient digital twin (SPDT).

TL;DR: The Sarconnector combines a SQL/R back end with a Shiny-based front end to capture CROMS, PROMS, PREMS, quality indicators, and health economics data across sarcoma centers. Its key innovation is meta-level analysis, identifying systematic practice differences across institutions rather than just describing individual patients. The platform's end-state is a sarcoma patient digital twin (SPDT) built from accumulated real-world-time data.
Pages 6-8
Striking Differences in Biopsy and Chemotherapy Rates Between Two Comparable Centers

The most clinically consequential finding of this study is that two large, high-volume sarcoma centers, both operating under established international guidelines and employing comparable multidisciplinary team structures, show substantial and statistically significant differences in core practice parameters. MDT/SB-A had approximately twice as many first-time presentations as MDT/SB-B over the same 15-month period, likely reflecting differences in referral networks and case complexity distribution, while follow-up presentations were comparable between the two centers.

Biopsy rates: At MDT/SB-A, 523 of 610 patients (85.7%) underwent biopsy, compared with 259 of 373 patients (69.4%) at MDT/SB-B. This 16.3 percentage-point difference was highly statistically significant (p < 0.0001). The authors partially attribute the higher biopsy rate at MDT/SB-A to a greater proportion of benign lesions being presented, requiring histological confirmation to rule out malignancy. However, this explanation does not fully account for the gap, and the difference may also reflect genuine variation in clinical judgment about when biopsy is indicated prior to surgical resection or multimodal therapy.

Chemotherapy rates: The discrepancy in chemotherapy use was even more dramatic. At MDT/SB-A, only 23 of 330 sarcoma patients (6.9%) received chemotherapy, compared with 83 of 304 sarcoma patients at MDT/SB-B (27.3%), a difference of 20.4 percentage points that was also highly significant (p < 0.0001). This fourfold difference in chemotherapy utilization between centers treating equivalent patient populations under the same guidelines points to fundamentally different institutional treatment philosophies, or to differences in patient case mix that the aggregate data cannot fully disentangle. Both the biopsy and chemotherapy differences were formally confirmed by Fisher's exact test.

The authors frame these findings as evidence for the necessity of benchmarking, not as a verdict on which center is practicing better. Whether one center is over-performing biopsies or under-using chemotherapy, or vice versa, cannot be determined without longitudinal outcome data linked to these practice variables. The Sarconnector is designed to accumulate exactly this outcome data over time, eventually allowing the platform to associate practice patterns with patient outcomes and costs.

TL;DR: Despite both operating under equivalent international guidelines, MDT/SB-A biopsied 85.7% of patients versus 69.4% at MDT/SB-B (p < 0.0001), and used chemotherapy in only 6.9% of sarcoma cases versus 27.3% at MDT/SB-B (p < 0.0001). These statistically significant fourfold differences in chemotherapy use confirm that guideline adherence alone does not produce uniform care, making outcome-linked benchmarking essential.
Pages 8-10
Real-Time Visualization, Subgroup Analysis, and Automated Statistical Testing

A practical strength of the Sarconnector is its interactive front-end visualization layer, which allows clinicians and researchers to explore the data without requiring statistical programming skills. The platform supports side-by-side comparative visualization of any parameter across MDT/SBs, filterable by time period, anatomical tumor location, tumor histological dignity (benign, borderline, or malignant), or therapy type. The system automatically links each visualization to the underlying raw data, enabling seamless drill-down from aggregate trends to individual case records for more detailed analysis.

Automated statistical testing: The statistical module embedded in the Sarconnector automates the selection and execution of appropriate tests for any chosen comparison. The workflow begins with a Shapiro-Wilk normality test and Levene's test for equality of variances to determine distributional assumptions, supplemented by visual inspection of normal Q-Q plots and histograms. Based on these preliminary assessments, the system selects and performs the appropriate inferential test, whether parametric (t-test) or non-parametric, and generates summary statistics including means, standard deviations, and Kaplan-Meier survival estimates.

Advanced analytical capabilities: Beyond standard descriptive and comparative statistics, the Sarconnector supports advanced techniques including Cox proportional hazard regression and competing risk analysis, enabling longitudinal survival analysis as outcome data accumulates. Figures generated through this module are formatted to publication-ready standards, meaning they can be used directly in scientific manuscripts without post-processing. The chemotherapy and biopsy comparisons reported in this paper were produced using this automated statistical tool, serving as a concrete demonstration of the platform's analytical output.

The platform is designed to support the full spectrum of sarcoma research workflows: from generating instant summary statistics during MDT/SB meetings for clinical decision support, to producing sophisticated multivariate survival models as longitudinal follow-up data accumulates over years. This dual utility as both a clinical tool and a research infrastructure is central to its design philosophy.

TL;DR: The Sarconnector's front end allows interactive filtering and side-by-side MDT/SB visualization by tumor type, therapy, and time period. Its automated statistical module applies Shapiro-Wilk, Levene's test, t-tests, Cox regression, and competing risk analysis, and generates publication-ready figures. Advanced outputs include Kaplan-Meier estimates and multivariate survival models as longitudinal outcome data accumulates.
Pages 10-12
From Benchmarking to the Sarcoma Patient Digital Twin

The study positions the Sarconnector within a broader strategic vision for sarcoma care transformation, articulated through two interconnected frameworks. The first is value-based healthcare (VBHC), which holds that healthcare should be measured and reimbursed based on outcomes achieved relative to resources consumed. For this to function in sarcoma, outcome metrics must be granular (covering not just survival but functional outcomes, complications, and patient experience), longitudinal (tracked across the full care cycle from diagnosis through years of follow-up), and cost-informed (linked to resource utilization data). The Sarconnector's architecture, with its five data dimensions including health economics, is designed to meet all three requirements simultaneously.

Integrated practice units: The paper draws on Porter's concept of the integrated practice unit (IPU) as the organizational structure best suited to deliver VBHC. In an IPU, all specialists involved in treating a specific medical condition are co-organized around the patient rather than around separate departments. For sarcoma, the MDT/SB is already a partial implementation of the IPU concept. The Sarconnector extends this by digitally integrating the data flows that the MDT/SB generates, enabling the IPU to operate not just on individual patient data but on population-level performance intelligence derived from its entire patient registry.

The sarcoma patient digital twin: The paper's most forward-looking concept is the sarcoma patient digital twin (SPDT). A digital twin is a virtual computational model that mirrors a real-world entity in real time, continuously updated by incoming data and capable of simulating future states. Applied to sarcoma patients, the SPDT would integrate each patient's clinical trajectory, imaging biomarkers, treatment decisions, outcomes, and patient-reported experiences into a dynamic model. Over time, as the platform accumulates data from hundreds or thousands of patients, predictive AI and ML models can be trained on this rich RWTD to forecast individual patient outcomes, identify high-risk patients early, and support personalized treatment decisions. The authors describe this not as a distant aspiration but as the intended endpoint of the Sarconnector's longitudinal data accumulation strategy.

The VBHC and SPDT visions share a critical dependency: both require the structured, harmonized, interoperable data that the Sarconnector is designed to generate. Without a consistent data schema applied prospectively across institutions, neither meaningful benchmarking nor reliable predictive modeling is achievable. This paper's contribution is therefore not primarily its immediate clinical findings, but its establishment of the infrastructure and framework for both.

TL;DR: The Sarconnector operationalizes value-based healthcare by capturing outcomes, costs, CROMS, PROMS, and PREMS simultaneously across sarcoma IPUs. Its long-term goal is the sarcoma patient digital twin (SPDT): a dynamically updated computational model of each patient's care trajectory, trained on accumulated RWTD and capable of predicting outcomes and supporting personalized treatment decisions via AI/ML.
Pages 12-14
Current Constraints: Data Availability, Causality, and Generalizability

Data completeness and accuracy: The reliability of benchmarking conclusions from the Sarconnector depends entirely on the quality and completeness of the data entered. Missing or inaccurate data can bias comparisons between MDT/SBs in ways that are difficult to detect or correct after the fact. In clinical practice, data entry burden is a persistent challenge: clinicians are already time-pressured, and structured data entry platforms compete with other clinical documentation responsibilities. The authors acknowledge this as a real-world implementation constraint, though the platform's design attempts to minimize the data entry burden through a "minimal dataset" approach that requests only the most diagnostically and prognostically critical parameters for each discipline.

Representativeness: The two MDT/SBs included in this study are both large, high-volume, tertiary academic referral centers. While they qualify as internationally representative under the criteria used, they may not reflect the full diversity of sarcoma care delivery across community hospitals, lower-volume centers, or healthcare systems with different resource levels and guideline implementation. The differences observed between these two expert centers may actually underestimate the variation that would be found if lower-volume community oncology practices were included in the benchmark comparison.

Causality versus association: The Sarconnector identifies associations between care practices and institutional characteristics, not causal relationships. The statistically significant differences in biopsy and chemotherapy rates between MDT/SB-A and MDT/SB-B demonstrate that real variation exists, but they do not establish which center's approach is superior, whether the variation is appropriate given unmeasured differences in patient characteristics, or what the downstream outcome consequences are. Confounding factors such as differing patient referral patterns, surgeon experience profiles, histological subtype distributions, and local resource availability could partially explain the observed differences without implying a quality deficit at either center.

Scalability and adoption: The platform was designed and implemented by its creators at two cooperating institutions. Scaling this approach to regional, national, or international levels introduces substantial technical, organizational, and regulatory challenges. Data governance frameworks differ across jurisdictions, EHR systems use different data formats and coding standards, and patient privacy regulations (such as GDPR in Europe) constrain cross-border data exchange. Federated learning approaches, in which models are trained across institutions without centralizing raw patient data, are identified as a promising mitigation strategy, but have not yet been implemented in the Sarconnector as reported in this paper.

TL;DR: Key limitations are data completeness (missing entries can bias benchmarks), restricted representativeness (only two high-volume tertiary centers), inability to establish causality from the observed biopsy and chemotherapy differences, and scalability challenges from heterogeneous EHR systems and cross-jurisdictional data governance requirements. Federated learning is noted as a future direction, not yet implemented.
Pages 14-16
Toward Outcome-Linked Benchmarks, Predictive AI, and Sustainable Sarcoma Care

Linking process metrics to outcomes: The immediate next phase for the Sarconnector is the accumulation of longitudinal outcome data linked to the practice-level variables already captured. With sufficient follow-up, the platform will be able to answer whether centers with higher chemotherapy utilization achieve better disease-free or overall survival for comparable patient populations, whether higher biopsy rates are associated with improved diagnostic accuracy or treatment appropriateness, and whether PROMS and PREMS scores correlate with clinical outcome measures or with institutional practice patterns. These outcome linkages are essential for transforming the current descriptive benchmark into a prescriptive tool that can recommend best-practice parameters with evidence-based justification.

Predictive AI and ML on RWTD: As the Sarconnector's longitudinal dataset grows, the platform is intended to support supervised machine-learning models for outcome prediction. These models would be trained on the combination of clinical variables, treatment parameters, CROMS, PROMS, imaging data, and health economics dimensions to predict patient-specific probabilities of recurrence, survival, complication risk, and treatment response. The unique strength of RWTD over clinical trial data for this purpose is its capture of contextual, systemic, and temporal factors, including the sequential trajectory of decisions made over the full care cycle, rather than only the controlled snapshots that trial data provide.

Expanding to multi-institutional and international benchmarking: The authors envision scaling the Sarconnector from its current two-center proof-of-concept to regional and international sarcoma networks. The international cancer benchmarking partnership (ICBP) provides a precedent for cross-national cancer outcome comparison, but focuses on survival metrics alone. The Sarconnector aims to benchmark the full pathway of care delivery, from MDT/SB practice patterns through surgical complexity, radiation planning, and patient experience, across diverse healthcare ecosystems. This would enable identification of best-practice centers globally and facilitate knowledge transfer to lower-performing institutions.

Integration with FAIR data principles and federated learning: Future development of the platform will incorporate FAIR data principles (Findable, Accessible, Interoperable, Reusable) to ensure that the RWTD it generates meets standards for scientific reuse and cross-platform exchange. Federated learning architectures, in which AI models are trained collaboratively across sites without pooling raw patient data at a central server, are identified as essential for maintaining patient privacy while enabling large-scale model training on the rare and heterogeneous sarcoma patient populations distributed across international institutions.

TL;DR: Future directions include linking practice metrics to long-term outcomes, training supervised ML models on accumulated RWTD for individualized outcome prediction, expanding benchmarking to international sarcoma networks beyond the ICBP model, and implementing FAIR data principles alongside federated learning to enable privacy-preserving multi-institutional AI model training across rare sarcoma subtypes.