Autologous CD19-directed chimeric antigen receptor (CAR) T-cell therapy, including tisagenlecleucel (CTL019), has become a standard salvage option for adult patients with relapsed or refractory aggressive B-cell lymphomas. In pivotal trials, complete response rates range from approximately 40% to 58%, meaning a substantial proportion of patients do not achieve durable remissions despite undergoing a complex, expensive, and physically demanding treatment. The inability to predict who will respond before the cells are infused means clinicians cannot effectively triage patients toward CAR-T versus alternative strategies, nor can they counsel patients on realistic expectations.
Limitations of existing prognostic tools: The International Prognostic Index (IPI) is the standard clinical risk stratification instrument for diffuse large B-cell lymphoma (DLBCL), incorporating age, stage, LDH, performance status, and number of extranodal sites. However, as this study demonstrates, the IPI performs poorly at predicting CAR-T response, with accuracy of only 0.54, sensitivity of 0.38, and specificity of 0.61 in this cohort. This effectively makes IPI-based prediction barely better than chance in this context. Pathology-based biomarkers require invasive re-biopsy and sample only individual lesions, missing the spatial heterogeneity of disease distributed across multiple lymph node stations throughout the body.
The image-based rationale: Pre-treatment diagnostic CT and PET/CT scans are already obtained on every lymphoma patient undergoing restaging before CAR-T infusion. These images contain rich spatial and morphological information about tumor burden, lesion characteristics, and inter-lesion heterogeneity that is not captured by clinical scoring systems. Deep learning (DL) applied to pre-treatment images could, in principle, extract prognostic features invisible to the human eye and provide a noninvasive, whole-body assessment of disease character before any treatment is initiated.
This 2023 study from the University of Pennsylvania Medical Image Processing Group and Lymphoma Program (Abramson Cancer Center), funded by NIH grant R01 CA255748, sets out to test the feasibility of this approach using pre-treatment diagnostic CT (dCT), low-dose CT from PET/CT (lCT), and FDG-PET images in 39 patients with B-cell lymphomas treated with CD19-directed CAR T-cells.
The study cohort was derived from a clinical research protocol (ClinicalTrials.gov NCT02030834) at the University of Pennsylvania evaluating tisagenlecleucel for relapsed or refractory DLBCL and follicular lymphoma (FL). The 39-patient cohort comprised 26 patients with DLBCL (20 male, 6 female; median age 57 years, range 28-74) and 13 patients with FL (7 male, 6 female; median age 62 years, range 43-72). All patients had pre-treatment diagnostic CT and PET/CT imaging of the neck, chest, abdomen, and pelvis acquired as part of clinical care.
Lesion identification and ground truth: An expert radiologist identified all individual lymph node disease sites on pre-treatment images and established ground truth lesion-level treatment responses by comparing pre-treatment and post-treatment imaging acquired a mean of 94.0 plus or minus 33.2 days after baseline. Lesion response was defined as interval decrease in size or metabolic activity, or interval resolution. Non-response was defined as lack of change or interval increase in size or metabolic activity. Extranodal lesions, splenic lesions, and Waldeyer's ring lesions were excluded due to small numbers at those sites.
Dataset scale: In total, 770 lymph node lesions were analyzed across three imaging modalities: 402 lesions from dCT, 214 from lCT, and 154 from FDG-PET. This lesion-level dataset, while relatively small at the patient level (39 patients), provided a meaningful volume of individual training instances for the deep learning models. Patient-level response was additionally defined using 12-month post-treatment scans as the reference standard, providing a clinically relevant long-term outcome measure.
IPI risk factor distribution: Among the 26 DLBCL patients, IPI risk groups were distributed as follows: 8 patients had 0-1 risk factors (low risk), 11 had 2 factors (low-intermediate risk), 4 had 3 factors (high-intermediate risk), and 3 had 4-5 factors (high risk). The responder/non-responder split for patient-level prediction was 10 responders and 16 non-responders among DLBCL patients, reflecting the generally poor outcomes typical of relapsed/refractory disease.
The deep learning framework centered on transfer learning using AlexNet, a convolutional neural network originally trained on millions of non-medical images from the ImageNet dataset. AlexNet was chosen for its relatively simple architecture (5 convolutional layers) that facilitates retraining with limited medical imaging datasets. The pre-trained AlexNet was adapted for binary classification (responder vs. non-responder) by replacing the final three layers with a fully connected layer, a Softmax layer, and a binary classification output layer. This approach leverages the low-level visual features learned from natural images while adapting the decision-making layers to the lymphoma response prediction task.
Five input scenarios: The study systematically evaluated five distinct approaches to representing each lymph node lesion for DL input. These were: (1) a single VOI-restricted slice through the lesion mid-portion (1 VOI-slice), (2) three contiguous VOI-restricted slices (3 VOI-slices), (3) a single whole-image axial slice through the lesion (1 whole-slice), (4) three contiguous whole-image slices (3 whole-slices), and (5) combined VOI-restricted and whole-image slices in two channels of one input sample (combined-slices). VOI-restricted inputs isolate the immediate lesion region, whereas whole-slice inputs provide the full anatomical context of the body cross-section.
Incremental learning: For the lCT and PET modalities, incremental learning was also evaluated, in which the dCT transfer learning model was used as a starting point and fine-tuned with lCT or PET training samples. This allowed the study to examine whether knowledge learned from the higher-quality diagnostic CT could be transferred to PET/CT-derived images.
Experimental scale and cross-validation: A total of 3,040 experiments were conducted: 2,400 using transfer learning and 640 using incremental learning. Each experiment was repeated 10 times with different randomized 6:2:2 (training/validation/testing) dataset splits to improve statistical reliability. Data augmentation was applied to training sets to mitigate overfitting. Statistical comparisons used two-sided t-tests with a significance threshold of p less than 0.05. The CAVASS software platform was used for placing 3D rectangular boxes around lesions.
The clearest finding from the lesion-level experiments was that whole-slice input scenarios consistently and significantly outperformed VOI-restricted inputs. For dCT, the 1 whole-slice scenario achieved accuracy 0.82 plus or minus 0.05, sensitivity 0.87 plus or minus 0.07, specificity 0.77 plus or minus 0.12, and AUC 0.91 plus or minus 0.03. In direct contrast, the 1 VOI-slice scenario yielded accuracy 0.68 plus or minus 0.05 and AUC 0.59 plus or minus 0.04, a performance level only marginally above chance. The improvement from adding full anatomical context (the entire axial cross-section) over using only the isolated lesion region was statistically highly significant (p less than 0.0001 for both accuracy and AUC comparisons).
Performance across modalities: The lCT and PET modalities achieved even higher accuracy on whole-slice inputs than dCT. For lCT, the 1 whole-slice scenario achieved accuracy 0.91 plus or minus 0.06, sensitivity 0.94 plus or minus 0.06, specificity 0.75 plus or minus 0.32, and AUC 0.92 plus or minus 0.08. For PET, accuracy was 0.87 plus or minus 0.06, sensitivity 0.90 plus or minus 0.06, specificity 0.77 plus or minus 0.19, and AUC 0.93 plus or minus 0.07. The accuracy of dCT whole-slice was statistically lower than lCT (p=0.002), though AUC and specificity did not differ significantly between modalities.
Adding more slices and combined inputs: Expanding from 1 whole-slice to 3 whole-slices did not produce significant performance gains. For dCT, AUC was 0.91 plus or minus 0.03 for 1 whole-slice versus 0.90 plus or minus 0.05 for 3 whole-slices (p=0.435). Similarly, the combined-slice input (VOI plus whole-slice in two channels) was not statistically different from the 1 whole-slice scenario. This suggests that a single axial slice containing the full anatomical cross-section captures most of the prognostically relevant information, and that additional slices or VOI information do not meaningfully improve the model.
Incremental versus transfer learning: For lCT and PET, there were no significant differences in accuracy or AUC between transfer learning (fine-tuned from ImageNet pre-training) and incremental learning (fine-tuned from the dCT model) using the 1 whole-slice input. This suggests the features learned from dCT do not provide a meaningful additional advantage over ImageNet initialization for these modalities.
Because individual patients have multiple lymph node lesions, translating lesion-level DL predictions into a single patient-level response classification required an aggregation step. The authors designed a rule-based reasoning framework with three options. The "All" rule classified a patient as a responder only if every individual lesion was predicted to respond, and as a non-responder if at least one lesion was predicted not to respond. The "Majority 60%" rule classified a patient as a responder if at least 60% of lesions were predicted to respond. The "Majority 70%" rule used a 70% threshold. This rule-based approach was explicitly chosen over machine learning at the patient level because the 39-patient cohort was too small to train and reliably validate a patient-level ML classifier.
Best-performing rule for dCT: For all 39 patients using dCT imaging, the Majority 60% rule achieved accuracy 0.79, sensitivity 0.83, and specificity 0.75. Among the 26 DLBCL patients specifically, the Majority 60% rule achieved accuracy 0.81, sensitivity 0.75, and specificity 0.88. This strong specificity means the model was particularly effective at correctly identifying non-responders, which is clinically valuable for avoiding futile CAR-T treatment in patients unlikely to benefit.
Comparison across rules: For dCT, the Majority 60% rule (accuracy 0.79) was statistically significantly better than the All rule (accuracy 0.61, p=0.027) but not significantly different from the Majority 70% rule (accuracy 0.71, p=0.38). The All rule's high sensitivity (1.00 for DLBCL patients, meaning it never missed a true responder) comes at the cost of very low specificity (0.64), classifying most non-responders incorrectly as responders. The Majority rules strike a more clinically useful balance.
lCT and PET patient-level performance: Patient-level prediction from lCT (Majority 60%: accuracy 0.65, sensitivity 0.60, specificity 0.75 across all patients) and PET (Majority 60%: accuracy 0.56, sensitivity 0.55, specificity 0.57) were lower than dCT and not statistically significantly different from dCT results, with p=0.80 and p=0.87 respectively. The diagnostic CT thus performed best for patient-level prediction despite lCT and PET achieving higher lesion-level AUC values.
The International Prognostic Index (IPI) is the established clinical standard for risk stratification in DLBCL, and the study compared its predictive performance against the DL rule-based approach in the 26 DLBCL patients (10 responders, 16 non-responders). For IPI-based prediction, patients were categorized as responders or non-responders using three different IPI risk factor thresholds: IPI less than or equal to 1, IPI less than or equal to 2, and IPI less than or equal to 3. The best IPI threshold (IPI less than or equal to 1) yielded accuracy 0.54, sensitivity 0.38, and specificity 0.61. The other thresholds performed even worse: IPI less than or equal to 2 had accuracy 0.42, and IPI less than or equal to 3 had accuracy 0.27 with specificity of 0.00.
Statistical significance of the improvement: The DL Majority 60% dCT approach (accuracy 0.81, sensitivity 0.75, specificity 0.88) outperformed the best IPI-based prediction (accuracy 0.54) with a p-value of 0.046. This statistically significant difference establishes that imaging-based deep learning prediction is not merely incrementally better but represents a fundamentally different level of prognostic information for CAR-T response. The DL model achieved a sensitivity of 0.75 compared to 0.38 for IPI, meaning it detected twice as many true responders, which is critical for ensuring that patients likely to benefit are not steered away from effective therapy.
Context from the literature: Prior work by Reinart et al. examined CT-based textural features and PET parameters in DLBCL patients undergoing CAR-T therapy. While statistically significant differences in whole-body metabolic tumor volume, total lesion glycolysis, and CT texture properties were found between complete responders and partial responders at baseline, no prediction analysis using a separate test set was performed in that study. The work by Galaznik et al. and Biccler et al. used regression and machine learning on clinical/pathological variables to predict outcomes in DLBCL treated with standard-of-care chemotherapy, with Biccler et al. reporting a concordance index of 0.756 for Danish and 0.744 for Swedish cohorts. The current study is distinct in applying end-to-end deep learning directly to pre-treatment imaging for CAR-T-specific response prediction.
The fact that whole-slice rather than lesion-restricted inputs produced the best performance suggests that the DL model is extracting prognostic information not only from the tumor itself but also from the surrounding tissue microenvironment and the spatial distribution of other structures visible in the cross-section, information that no clinical scoring system attempts to capture.
Small patient sample: The most consequential limitation of this study is the cohort size of 39 patients. While 770 individual lesions provided a reasonable volume of training instances for lesion-level DL, the patient-level analysis is based on 39 cases, which is far too small to apply machine learning at the patient level. The authors explicitly acknowledge this, stating that the small number of patients "precluded use of machine learning approaches for patient-level response prediction." The rule-based reasoning approach was adopted as a practical workaround but carries its own limitations since it does not learn optimal aggregation strategies from data and uses fixed thresholds (60% or 70%) that may not generalize to other cohorts.
Restriction to lymph node lesions: The analysis was limited to conventional nodal lymphoma sites, excluding extranodal lesions, splenic involvement, and Waldeyer's ring lesions due to insufficient numbers of such lesions in the cohort. This is a meaningful limitation since extranodal disease is common in DLBCL and may carry different response characteristics. Future larger studies should incorporate these lesion types to enable a more comprehensive whole-body response prediction.
Single institution and retrospective design: All patients were treated at the University of Pennsylvania on a single clinical protocol, and the study was retrospective. This creates the potential for spectrum bias and limits the generalizability of the model to other institutions with different patient demographics, imaging protocols, scanner vendors, and clinical practices. External validation at independent centers is a critical prerequisite before clinical translation.
AlexNet architecture: The authors used AlexNet, a relatively simple and older CNN architecture, as the backbone. They explicitly note that the same framework could be reconfigured with more recent and performant architectures such as VGG or ResNet. Given that more modern architectures consistently outperform AlexNet on computer vision benchmarks, it is likely that accuracy and AUC could be further improved with updated backbone networks, though the current study establishes the feasibility proof of concept.
The clinical potential demonstrated by this study is significant: if a DL model applied to standard pre-treatment CT scans can predict with 81% accuracy which DLBCL patients will or will not respond to CAR-T therapy at 12 months, this information could directly influence clinical decision-making. Patients predicted to be non-responders might be preferentially enrolled in clinical trials testing novel combination strategies or alternative salvage therapies. Patients predicted to respond could proceed to CAR-T with greater confidence. The noninvasive nature of the approach, relying only on images already acquired as part of standard care, is a key practical advantage over biomarker approaches requiring new biopsies.
The whole-slice finding and its implications: The consistent superiority of whole-slice over VOI-restricted inputs is a methodologically important finding with broad implications for medical image-based DL. It suggests that the prognostic signal is not contained solely within the tumor itself but is distributed across the entire body cross-section visible in the axial slice. This could reflect relationships between tumor morphology and adjacent normal tissue, lymph node architecture in surrounding regions, body composition features, or other contextual features. Understanding which image regions the model is attending to (via interpretability methods such as class activation mapping or Grad-CAM) would be an important direction for follow-up work.
Extending to larger and more diverse cohorts: The authors identify several clear paths forward, including expanding the dataset to larger patient numbers from multiple institutions, incorporating extranodal and splenic lesions, and testing more contemporary DL architectures. Federated learning across institutions treating CAR-T patients could enable training on much larger datasets while preserving patient privacy. Integration with other pre-infusion biomarkers, such as circulating tumor DNA (ctDNA), CAR-T cell product characteristics, inflammatory cytokine levels, and tumor mutational burden, could further improve prediction through multimodal fusion models.
Broader applicability: While this study focused specifically on CAR-T therapy prediction, the same imaging-based DL framework is in principle applicable to predicting response to other lymphoma treatments, including standard R-CHOP chemoimmunotherapy, bispecific antibody therapies, and antibody-drug conjugates. The availability of pre-treatment PET/CT images as a universal staging tool in lymphoma creates a large potential reservoir of training data for such prediction models once ground truth response labels are systematically collected. This study's contribution is to establish the proof of concept and provide performance benchmarks against which future, larger-scale imaging AI tools for CAR-T and lymphoma response prediction can be compared.