CPU-Index: Concordance-Based Predictive Uncertainty for Improved Lung Cancer Screening Specificity

Artif Intell Med 2025 AI 7 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Reducing False Positives in Lung Cancer CT Screening

The False Positive Problem in Lung Cancer Screening Low-dose CT (LDCT) screening can catch lung cancer early when it is most curable, but it produces a high rate of false positive results - CT findings that look suspicious but turn out not to be cancer. False positives lead to unnecessary follow-up scans, invasive biopsies, patient anxiety, and healthcare costs. Improving screening specificity (reducing false positives) without sacrificing sensitivity is a major clinical goal.

Current Prediction Models and Their Limitations Existing AI models for lung cancer screening predict risk based on radiomic features from CT scans and clinical data, but they do not account for uncertainty in their predictions. A model might give a high-risk score based on features that are uncertain or noisy, leading to confident-looking predictions that are actually unreliable.

The CPU-Index Innovation The Concordance-based Predictive Uncertainty (CPU)-Index is a novel uncertainty quantification framework. Rather than just outputting a risk score, the CPU-Index measures how consistent the model's prediction is with predictions made for similar patients. High concordance between the individual patient and their closest matches gives confidence in the prediction; low concordance flags uncertainty.

Study Scale and Design The study analyzed 1,767 patients from an initial cohort of 3,326 low-dose CT screening participants, using radiomic features (172 features per patient) combined with electronic health record (EHR) data. The CPU-Index was evaluated against a baseline prediction model to measure improvement in AUC and false positive rate.

TL;DR: The CPU-Index is a new framework that quantifies uncertainty in lung cancer risk predictions by checking whether a model's prediction is consistent with similar patients' predictions, reducing false positive screening rates.
Pages 2-4
NMTLR Model and Positional Encoding for Feature Fusion

Base Prediction Model: NMTLR The Non-linear Monotone Time-to-Event Learning via Regression (NMTLR) model was used as the baseline survival prediction framework. Unlike simple classification models, NMTLR is a time-to-event (survival) model that predicts when an event (cancer diagnosis) will occur rather than just whether it will occur. This captures the temporal dimension of cancer screening data.

Radiomic Feature Extraction 172 radiomic features were extracted from each patient's LDCT scan, capturing first-order statistics, texture (GLCM, GLRLM, GLSZM matrices), shape characteristics, and wavelet transform features from detected lung nodules. These features encoded subtle imaging patterns related to nodule morphology and growth characteristics.

Positional Encoding for Multimodal Fusion A key technical innovation was positional encoding to fuse radiomic imaging features with structured EHR data (age, smoking history, pack-years, spirometry values, etc.). Positional encoding is a technique borrowed from transformer neural networks that encodes the relative position or type of each feature, enabling heterogeneous data types from different domains to be combined meaningfully.

Feature Harmonization Across Sites Since the dataset combined patients from multiple imaging centers with different CT scanners and protocols, ComBat harmonization was applied to the radiomic features to remove scanner-specific batch effects while preserving biological variation. This step is essential for radiomic models applied across multi-site datasets.

TL;DR: The NMTLR survival model combined 172 CT radiomic features with EHR data using positional encoding for multimodal fusion, with ComBat harmonization to correct for differences between imaging sites.
Pages 4-6
How the CPU-Index Measures Prediction Uncertainty

The Core Concept: Concordance Between Similarity and Prediction The CPU-Index is built on an elegant insight: if the model's prediction for a patient is reliable, then similar patients (those with similar radiomic and clinical profiles) should receive similar predictions. If a patient receives a very different prediction from their closest neighbors, this inconsistency signals that the prediction is unreliable.

Subgroup Similarity Measurement For each patient, a subgroup of the most similar patients was identified using nearest-neighbor algorithms applied to the radiomic-EHR feature space. Similarity was measured using distance metrics in the high-dimensional feature space, with cross-validated selection of neighborhood size to optimize the uncertainty estimate.

Concordance Index Calculation The concordance between the patient's risk prediction and the predictions of their subgroup was quantified using a concordance index. High concordance (patient's prediction is consistent with similar patients) yields a low CPU-Index (high confidence). Low concordance (patient's prediction diverges from similar patients) yields a high CPU-Index (high uncertainty).

1000-Permutation Testing for Validity To confirm that the CPU-Index measures genuine prediction uncertainty rather than random variation, 1,000 permutation tests were performed. Predicted labels were randomly shuffled and CPU-Index distributions were calculated, confirming that the real CPU-Index distributions were significantly different from chance, validating its statistical meaningfulness.

TL;DR: The CPU-Index measures how consistent a patient's risk prediction is with those of similar patients; high consistency means low uncertainty, while prediction divergence from similar patients signals unreliable results.
Pages 6-8
CPU-Index Substantially Improves Screening Performance

AUC Improvement Adding the CPU-Index to the baseline NMTLR model raised the AUC from 0.81 to 0.89. This 0.08-point AUC improvement is clinically meaningful, representing a substantially better ability to discriminate true lung cancer from benign findings in screening CT scans.

False Positive Rate Reduction Most practically important for patients, the CPU-Index reduced the false positive rate from 0.41 to 0.30 at the same sensitivity threshold. This means 11 fewer false positive results per 100 patients screened - translating to thousands of avoided unnecessary follow-up procedures in a large screening program.

High-Uncertainty Cases Analysis Patients with high CPU-Index values (high prediction uncertainty) had substantially different characteristics from low-uncertainty patients. High-uncertainty cases tended to have unusual radiomic profiles that diverged from the training distribution - exactly the cases where model predictions are least reliable and human expert judgment is most needed.

Clinical Workflow Integration The CPU-Index enables a two-tier clinical approach: patients with high confidence predictions (low CPU-Index) can be triaged based on the risk score alone, while patients with uncertain predictions (high CPU-Index) are flagged for additional clinical review, specialist consultation, or more detailed imaging protocols.

TL;DR: The CPU-Index raised AUC from 0.81 to 0.89 and reduced false positive rates from 41% to 30%, meaning 11 fewer false positive results per 100 screened patients without missing true cancer cases.
Pages 8-9
Understanding Which Cases Drive Uncertainty

Radiomic Feature Contributions to Uncertainty Analysis of high-uncertainty cases revealed specific radiomic features that frequently drove prediction inconsistency. Certain texture features related to nodule density heterogeneity and calcification patterns were most associated with high CPU-Index values, suggesting these features correspond to nodule subtypes that current models handle less reliably.

Patient Characteristics of Uncertain Cases High-uncertainty cases were more common among patients with unusual clinical profiles: very young or very old ages, atypical smoking histories, or specific comorbidities not well-represented in the training cohort. This confirms that uncertainty is highest where the training data had less coverage.

Comparison with Existing Uncertainty Methods The CPU-Index was compared against other uncertainty quantification approaches including Monte Carlo Dropout and deep ensembles. CPU-Index outperformed these alternatives in AUC improvement while requiring less computational overhead, making it more practical for large-scale screening deployment.

Subgroup Analysis Across Patient Demographics CPU-Index benefits were consistent across sex, age groups, and smoking status subgroups, suggesting the uncertainty framework is broadly applicable rather than beneficial only for specific patient types. Performance improvements were most pronounced in intermediate-risk patients where clinical decision-making is most challenging.

TL;DR: High uncertainty cases clustered around unusual nodule textures and atypical patient profiles underrepresented in training data, confirming that the CPU-Index correctly identifies where predictions are most likely to be unreliable.
Pages 9-10
Transforming Lung Cancer Screening Decision-Making

Actionable Two-Tier Triage System The CPU-Index enables a practical clinical triage: high confidence + high risk = prioritize for immediate further workup; high confidence + low risk = routine follow-up; uncertain prediction = expert review regardless of risk score. This nuanced approach is more clinically appropriate than treating all AI predictions equally.

Patient Counseling Applications Presenting uncertainty estimates to patients alongside risk scores enables more honest and nuanced informed consent discussions. Rather than stating a patient has a '30% cancer risk,' clinicians can explain that the model is confident vs. uncertain in that estimate, setting appropriate expectations and avoiding both false reassurance and unnecessary anxiety.

Quality Assurance for AI Deployment In clinical AI deployment, models frequently encounter patient cases that differ from training data. The CPU-Index provides a built-in quality check that automatically flags when the model is operating outside its reliable performance range, creating a safety mechanism for AI-assisted screening programs.

Implications for Screening Guidelines Current lung cancer screening guidelines (e.g., US Preventive Services Task Force) use simple smoking history thresholds to define screening eligibility. An AI system combining risk prediction with uncertainty quantification could enable more personalized, risk-based screening recommendations that better balance benefit (early detection) against harm (false positives).

TL;DR: The CPU-Index enables a practical two-tier triage where high-confidence predictions drive automated decisions and uncertain predictions trigger expert review, creating a safety mechanism for AI-assisted screening.
Pages 10-11
Limitations and Directions for Future Research

Single Dataset Validation Despite internal validation with 1,000 permutation tests, the CPU-Index framework was validated in a single dataset of 1,767 patients. External validation in independent screening cohorts from different countries and healthcare systems is needed before recommending clinical deployment.

Radiomic Reproducibility Challenges Radiomic features remain susceptible to CT acquisition variability. Despite ComBat harmonization, residual scanner-specific effects may influence the feature space used for nearest-neighbor similarity calculations, potentially affecting CPU-Index accuracy in settings with highly heterogeneous scanning protocols.

Expanding Beyond Lung Cancer Screening The CPU-Index framework is theoretically applicable to any AI prediction problem where uncertainty quantification would improve clinical decision-making. Future work should evaluate its value in other oncology screening contexts (colon cancer, breast cancer), diagnostic radiology, and treatment response prediction.

Integration with Emerging Technologies Future development should integrate the CPU-Index with foundation models trained on massive multi-center CT datasets, which would provide richer feature representations and larger, more diverse neighbor pools for concordance calculation. This could further improve uncertainty estimation and reduce the need for site-specific harmonization.

TL;DR: External validation across diverse screening populations is needed, and the CPU-Index framework has broad potential applicability beyond lung cancer screening to any AI prediction problem in oncology.
Citation: Open Access, 2025. Available at: PMC12139530.