Intentional deep overfit learning for patient-specific dose predictions in adaptive radiotherapy

Med Phys 2023 Treatment 6 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-3
The Problem of Anatomy That Changes During Radiation Treatment

Radiation therapy works by delivering precisely targeted high doses of radiation to a tumor while protecting the surrounding healthy tissue. Each patient's treatment plan is designed based on a CT scan taken before treatment begins, specifying exactly how the radiation beams should be shaped and directed. But the human body is not static: tumors shrink as they respond to treatment, patients lose weight, and organs shift position from day to day. The plan designed on day one may no longer be optimal by week three.

These changes create two risks. First, if the tumor has shifted, radiation may miss part of it -- undertreating the cancer. Second, if an organ at risk (such as the spinal cord, parotid gland, or bowel) has moved closer to the beam path, it may receive more radiation than planned, causing toxicity. This problem of intra-patient anatomical variation throughout a course of radiation is well documented, particularly in head and neck cancer where tumor shrinkage and weight loss are rapid and predictable.

Adaptive radiation therapy (ART) was developed specifically to address this problem. Rather than using a fixed plan for all 30 to 35 treatment sessions, ART adjusts the plan based on daily imaging, re-optimizing the dose distribution to account for how the patient's anatomy has changed. Modern online ART machines such as Varian's Ethos, Elekta's Unity, and ViewRay's MRIdian can acquire a new scan and recalculate the dose on the day of each treatment session.

However, re-optimizing a treatment plan still requires significant manual work from radiation oncologists and treatment planners, who must review and adjust optimization criteria for each session. This bottleneck limits how often adaptation can practically occur. This study from UT Southwestern Medical Center proposes a deep learning approach that pre-learns each individual patient's dosimetric preferences and automates dose prediction for their subsequent adaptive sessions, dramatically reducing the required manual effort.

TL;DR: Radiation therapy plans become suboptimal as patient anatomy changes during treatment; adaptive radiotherapy corrects this but requires time-consuming daily re-optimization that AI could streamline.
Pages 3-5
Population Models vs. Patient-Specific Models: The IDOL Approach

Conventional deep learning dose prediction models are trained on data from many patients and produce a single population model that learns average relationships between anatomy and optimal dose distribution. These models work reasonably well because different patients with similar cancer types often require similar dosimetric trade-offs. However, they cannot account for the particular preferences and constraints that a specific physician used when creating a specific patient's plan.

The key innovation in this study is using a technique called Intentional Deep Overfit Learning (IDOL), originally proposed by Chun et al. The approach deliberately does what is normally considered a cardinal sin in machine learning: it overfits a model to a single patient's data. Starting from the population model as a foundation, the model is fine-tuned using only the single initial treatment plan created for one specific patient. The result is a model that deeply memorizes that patient's dosimetric intent.

This is called intentional overfitting because the goal is not to generalize to other patients -- it is to become maximally accurate at predicting what the right dose distribution should be for this one patient, given their anatomy on any given day. The physician's original plan encodes their clinical judgment about how to balance tumor coverage against organ protection for that specific patient, and IDOL transfers this judgment into the model through fine-tuning.

The deep learning architecture used was the Hierarchically Densely Connected U-Net (HD U-Net), a 3D convolutional neural network that takes the planning target volume (PTV) and 44 organs at risk as input channels and outputs a predicted 3D dose distribution. Inputs were converted to 5mm voxel spacing, and random patch selection plus rotations and flips were used for data augmentation during training.

TL;DR: IDOL fine-tunes a population dose model using only one patient's initial treatment plan, deliberately overfitting to capture their physician's dosimetric intent for use in subsequent adaptive sessions.
Pages 5-6
Study Design: 43 Patients, Two Models, 26 Adaptive Test Sessions

The study used 43 head and neck cancer patients treated on the Varian Ethos adaptive radiotherapy machine. Each patient had an initial radiation therapy plan and a variable number of adaptive plans generated over their treatment course. In total, the dataset contained 43 initial plans and 88 adaptive plans.

The 43 patients were split into two groups. Thirty-three patients and all their plans (85 total: 30 initial plus 55 adaptive) were used to train the adaptive population (AP) model, which learns general dose-anatomy relationships across many patients. The remaining 10 patients were used differently: their 10 initial plans trained 10 individual fine-tuned patient-specific (FT-PS) models, while their 26 adaptive plans were completely withheld from all training as the held-out test set.

For each of the 26 test adaptive sessions, dose was predicted by both the population AP model and the patient-specific FT-PS model for the corresponding patient. Both predictions were compared against the ground truth dose distribution as originally calculated during the patient's actual treatment. Performance was measured using Mean Absolute Percent Error (MAPE), normalized to the patient's prescription dose, within the volume receiving at least 10 percent of the prescription dose.

Fine-tuning of FT-PS models was run for 5,000 iterations on a single NVIDIA V100 GPU, taking approximately 1 hour per patient. Once trained, generating a dose prediction for a new adaptive session takes under 1 minute. Population model training started from random initialization; fine-tuning started from the population model's saved weights.

TL;DR: Ten patient-specific models were fine-tuned from a shared population model using only each patient's initial plan, then tested on 26 withheld adaptive treatment sessions.
Page 7
Patient-Specific Models Cut Dose Prediction Error by 35 Percent

Across all 26 test adaptive sessions, the population model (AP) achieved an average MAPE of 5.759 percent, while the fine-tuned patient-specific models (FT-PS) achieved an average MAPE of 3.747 percent -- a reduction of 35 percent in prediction error. A paired t-test confirmed this difference was highly statistically significant, with a two-tailed p-value of 3.851 times ten to the negative ninth power (p less than 0.001) and a 95 percent confidence interval of negative 2.483 to negative 1.542, indicating the FT-PS model consistently outperformed the AP model across the patient cohort.

At the level of individual structures, statistically significant MAPE improvements with the FT-PS model were found in 19 out of 24 segmented structures. For maximum dose prediction (DMAX) per structure, significant improvements were seen in 13 out of 24 structures. These results demonstrate that patient-specific fine-tuning improves dose accuracy across most clinically relevant organs at risk, not just on average across the full treatment volume.

For the primary tumor target (PTV), there was no statistically significant difference between models for PTV coverage metrics including D98 percent, D99 percent, homogeneity index, and conformity index. Both models predicted equivalent coverage of the tumor target -- the key difference between them was in how accurately each modeled dose delivery to the surrounding normal tissues.

Dose distribution visualizations for an individual patient showed the FT-PS model more closely reproducing the spatial dose gradients at OAR boundaries as present in the ground truth, while the AP model showed visible discrepancies in regions near critical structures. This visual evidence reinforces the quantitative MAPE differences.

TL;DR: Patient-specific models reduced dose prediction error from 5.76 percent to 3.75 percent (35 percent improvement, p less than 0.001), with significant gains in 19 of 24 organs at risk.
Pages 8-9
Clinical Integration: From Daily Predictions to Automated Plan Optimization

The authors propose two specific clinical applications of FT-PS dose prediction models. The first is using the model's predicted dose distribution as the starting point for generating automated treatment planning optimization objectives. Rather than a planner manually setting objectives for each adaptive session, the FT-PS model would automatically generate objectives that reflect both the physician's original intent and the patient's current anatomy, reducing repetitive manual work without removing clinical oversight.

The second proposed application is an alert-based monitoring system that tracks FT-PS dose predictions session by session and flags significant changes. For example, if a patient's tumor shrinks away from a critical structure over the first few weeks of treatment, the system could alert the team that a plan modification to spare that structure is now possible. Conversely, if the system detects unexpected dose creep to an organ at risk, it could alert clinicians before a toxicity threshold is reached.

A key practical advantage is the training timeline. FT-PS model fine-tuning takes approximately 1 hour per patient, which can be completed in advance when the initial treatment plan is finalized, typically several days before treatment begins. Generating a dose prediction for a single new adaptive session then takes under 1 minute, fitting within the time constraints of online ART workflows where every minute the patient is on the treatment table matters.

The study also discusses an important conceptual question: whether the FT-PS model is truly patient-specific or more precisely initial-plan-specific. Because fine-tuning used only the initial plan, the model encodes the dosimetric constraints and priorities from that particular plan. The authors note this distinction may not matter practically: what is important is that the model accurately reflects the appropriate dose distribution for the remainder of that patient's treatment course.

TL;DR: Patient-specific models enable automated generation of adaptive planning objectives and dose-change alerts, taking only 1 hour to fine-tune and under 1 minute to predict per session in clinical workflows.
Page 9
Limitations and Future Steps Toward Routine Adaptive Radiotherapy

The study's main limitations relate to its small dataset. Only 43 patients were available because the Varian Ethos machine was relatively new at the time, limiting the pool of eligible patients. With only 10 patients used to test FT-PS models, generalizability of the results across a broader cancer population remains uncertain. The authors acknowledge that a larger population model trained on more patients could potentially narrow the gap with patient-specific models.

A fundamental challenge specific to IDOL training is the absence of a validation dataset. Normal machine learning practice uses a validation set to stop training at the optimal point; here, since only one plan per patient is available for training, training was stopped at an arbitrary cutoff of 5,000 iterations. The authors note that future work could use computer-generated deformed versions of the initial plan as synthetic validation data to determine a principled stopping point.

The models in this study used only the segmented structure masks and dose distributions as inputs, omitting information like the CT image itself and the beam geometry. Adding these data sources would give the model richer spatial context about each patient's anatomy and treatment configuration, likely improving prediction accuracy further.

Future work planned by the team includes incorporating additional patient-specific inputs (CT image, field geometry), performing continuous re-fine-tuning of the FT-PS model as new adaptive sessions occur (allowing the model to learn from each session), and formally integrating the system into the clinical ART workflow at UT Southwestern to measure its real-world impact on planning time, plan quality, and patient outcomes.

TL;DR: Limited by a small dataset and no validation set for FT-PS training, the study's findings nonetheless establish deep overfitting as a viable path to personalized adaptive radiotherapy, with future work aimed at clinical deployment.
Citation: Open Access, . Available at: PMC10530457.