Unmet Need: Lung cancer patients have complex and often unmet supportive care needs spanning psychological, physical, and social domains, yet systematic identification and prediction of these needs is rarely done.
Study Aim: This study assessed the supportive care needs of hospitalized lung cancer patients in China using the validated SCNS-SF34 questionnaire, and then built machine learning models to predict individual need levels.
Population: 486 hospitalized lung cancer patients were enrolled, representing a clinically realistic sample of patients with varying disease stages, treatment exposures, and demographic backgrounds.
Predictive Framework: Six machine learning algorithms were compared to identify the best model for predicting supportive care need scores, with the goal of enabling proactive, personalized support allocation.
SCNS-SF34 Instrument: The Supportive Care Needs Survey Short Form 34 assesses 34 items across five domains: psychological, health system/information, physical/daily living, patient care/support, and sexuality needs.
Response Format: Patients rate each item on a 5-point scale from 'no need' to 'high need,' generating domain-specific and total need scores that capture the multidimensional nature of cancer-related support requirements.
Validated for Cancer Patients: The SCNS-SF34 has been validated in multiple cancer populations and languages, providing a standardized, well-accepted measure for comparing needs across studies and patient groups.
Chinese Adaptation: The Chinese version of the SCNS-SF34 was used, culturally adapted for the study population to ensure linguistic and conceptual equivalence with the original instrument.
Six ML Algorithms Compared: The study evaluated logistic regression, support vector machine, decision tree, k-nearest neighbors, gradient boosting, and random forest to identify the most accurate predictor of supportive care need levels.
Predictor Variables: Input features included demographic characteristics (age, sex, education level, household income, marital status), clinical variables (tumor stage, pathological type, treatment status), and functional status.
Performance Metrics: Models were evaluated on mean absolute error (MAE), accuracy, F1 score, and AUC across cross-validation folds, providing a comprehensive comparison beyond simple accuracy.
Feature Importance Analysis: Variable importance scores were extracted from the best-performing model to identify which patient characteristics most strongly influenced predicted supportive care need levels.
Best Performer: Random Forest achieved the highest performance across all metrics: MAE 4.45, accuracy 88.42%, F1 score 87.49%, and AUC 0.9061, outperforming all five other algorithms.
AUC Interpretation: An AUC of 0.906 indicates excellent discriminative ability for identifying patients with high versus low supportive care needs, which is the actionable distinction for clinical care allocation.
Model Stability: Random forest's ensemble nature (combining many decision trees) likely contributed to its robustness, as it is less susceptible to overfitting on the moderate sample size of 486 patients.
Clinical Need Levels: The distribution of SCNS-SF34 scores across the 486 patients showed that many patients had moderate-to-high needs across multiple domains, confirming the importance of systematic assessment.
Education Level: Lower education was among the strongest predictors of higher supportive care needs, potentially reflecting reduced health literacy, limited coping resources, and greater reliance on healthcare providers for information.
Household Income: Lower household income was associated with higher needs, consistent with the financial toxicity of cancer care and limited access to private support services outside the hospital setting.
Age: Older patients showed distinct supportive care need profiles, with higher physical and care support needs but potentially different psychological need expressions compared to younger patients.
Tumor Stage and Pathological Type: Advanced disease stage and more aggressive histological subtypes were associated with higher overall need levels, reflecting the greater physical and psychological burden of advanced lung cancer.
Proactive Screening: Deploying the random forest model at hospital admission could automatically flag patients predicted to have high supportive care needs, enabling earlier referrals to social workers, counselors, or palliative care teams.
Resource Allocation: In healthcare systems with limited supportive care resources, risk stratification by predicted need level allows prioritization of intensive support to patients most likely to benefit.
Equity Considerations: The prominence of education and income as predictors highlights socioeconomic disparities in cancer support needs, informing policy decisions about targeted outreach to vulnerable populations.
Multidisciplinary Integration: Embedding need prediction into electronic health records can facilitate communication among oncology, nursing, social work, and palliative care teams to ensure coordinated, holistic patient support.
Single-Center Design: Data from one hospital in China may not generalize to other regions, health systems, or cultural contexts, limiting the external validity of the models.
Cross-Sectional Assessment: Supportive care needs were measured at a single time point, whereas needs evolve throughout the cancer treatment trajectory; longitudinal models would better capture this dynamic.
Chinese Healthcare Context: Healthcare system structure, family support norms, and patient-provider communication patterns in China may differ substantially from Western settings, affecting model transferability.
Future Directions: Prospective validation across multiple institutions, extension to longitudinal need prediction, and integration of quality-of-life outcome data would strengthen the clinical utility of this approach.