Exploring Supportive Care Needs of Lung Cancer Patients in China and Predicting with Machine Learning Models

Support Care Cancer 2025 AI 7 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Predicting Supportive Care Needs in Lung Cancer Patients

Unmet Need: Lung cancer patients have complex and often unmet supportive care needs spanning psychological, physical, and social domains, yet systematic identification and prediction of these needs is rarely done.

Study Aim: This study assessed the supportive care needs of hospitalized lung cancer patients in China using the validated SCNS-SF34 questionnaire, and then built machine learning models to predict individual need levels.

Population: 486 hospitalized lung cancer patients were enrolled, representing a clinically realistic sample of patients with varying disease stages, treatment exposures, and demographic backgrounds.

Predictive Framework: Six machine learning algorithms were compared to identify the best model for predicting supportive care need scores, with the goal of enabling proactive, personalized support allocation.

TL;DR: This study applied six ML models to predict supportive care needs in 486 lung cancer patients, with random forest achieving the best performance (AUC 0.906, accuracy 88.4%).
Pages 2-3
SCNS-SF34: Measuring Supportive Care Needs

SCNS-SF34 Instrument: The Supportive Care Needs Survey Short Form 34 assesses 34 items across five domains: psychological, health system/information, physical/daily living, patient care/support, and sexuality needs.

Response Format: Patients rate each item on a 5-point scale from 'no need' to 'high need,' generating domain-specific and total need scores that capture the multidimensional nature of cancer-related support requirements.

Validated for Cancer Patients: The SCNS-SF34 has been validated in multiple cancer populations and languages, providing a standardized, well-accepted measure for comparing needs across studies and patient groups.

Chinese Adaptation: The Chinese version of the SCNS-SF34 was used, culturally adapted for the study population to ensure linguistic and conceptual equivalence with the original instrument.

TL;DR: The SCNS-SF34 is a validated 34-item tool measuring five domains of supportive care need, used here in Chinese-language form to assess 486 hospitalized lung cancer patients.
Pages 3, 5
Machine Learning Model Development and Comparison

Six ML Algorithms Compared: The study evaluated logistic regression, support vector machine, decision tree, k-nearest neighbors, gradient boosting, and random forest to identify the most accurate predictor of supportive care need levels.

Predictor Variables: Input features included demographic characteristics (age, sex, education level, household income, marital status), clinical variables (tumor stage, pathological type, treatment status), and functional status.

Performance Metrics: Models were evaluated on mean absolute error (MAE), accuracy, F1 score, and AUC across cross-validation folds, providing a comprehensive comparison beyond simple accuracy.

Feature Importance Analysis: Variable importance scores were extracted from the best-performing model to identify which patient characteristics most strongly influenced predicted supportive care need levels.

TL;DR: Six ML algorithms were trained on demographic and clinical features to predict SCNS-SF34 scores, with comprehensive performance evaluation including MAE, accuracy, F1, and AUC.
Pages 6-7
Random Forest as the Best Predictive Model

Best Performer: Random Forest achieved the highest performance across all metrics: MAE 4.45, accuracy 88.42%, F1 score 87.49%, and AUC 0.9061, outperforming all five other algorithms.

AUC Interpretation: An AUC of 0.906 indicates excellent discriminative ability for identifying patients with high versus low supportive care needs, which is the actionable distinction for clinical care allocation.

Model Stability: Random forest's ensemble nature (combining many decision trees) likely contributed to its robustness, as it is less susceptible to overfitting on the moderate sample size of 486 patients.

Clinical Need Levels: The distribution of SCNS-SF34 scores across the 486 patients showed that many patients had moderate-to-high needs across multiple domains, confirming the importance of systematic assessment.

TL;DR: Random forest outperformed all other models with AUC 0.906 and 88.4% accuracy, making it the recommended algorithm for predicting supportive care needs in lung cancer patients.
Pages 7-8
Factors Most Strongly Associated with Supportive Care Needs

Education Level: Lower education was among the strongest predictors of higher supportive care needs, potentially reflecting reduced health literacy, limited coping resources, and greater reliance on healthcare providers for information.

Household Income: Lower household income was associated with higher needs, consistent with the financial toxicity of cancer care and limited access to private support services outside the hospital setting.

Age: Older patients showed distinct supportive care need profiles, with higher physical and care support needs but potentially different psychological need expressions compared to younger patients.

Tumor Stage and Pathological Type: Advanced disease stage and more aggressive histological subtypes were associated with higher overall need levels, reflecting the greater physical and psychological burden of advanced lung cancer.

TL;DR: Education, income, age, tumor stage, and pathological type were the strongest predictors of supportive care needs, reflecting both socioeconomic and disease-related determinants.
Pages 9-10
Clinical Application and Healthcare System Relevance

Proactive Screening: Deploying the random forest model at hospital admission could automatically flag patients predicted to have high supportive care needs, enabling earlier referrals to social workers, counselors, or palliative care teams.

Resource Allocation: In healthcare systems with limited supportive care resources, risk stratification by predicted need level allows prioritization of intensive support to patients most likely to benefit.

Equity Considerations: The prominence of education and income as predictors highlights socioeconomic disparities in cancer support needs, informing policy decisions about targeted outreach to vulnerable populations.

Multidisciplinary Integration: Embedding need prediction into electronic health records can facilitate communication among oncology, nursing, social work, and palliative care teams to ensure coordinated, holistic patient support.

TL;DR: ML-based need prediction at admission could enable proactive, equitable supportive care allocation in oncology settings, particularly for socioeconomically vulnerable patients.
Pages 11-12
Study Limitations and Future Research

Single-Center Design: Data from one hospital in China may not generalize to other regions, health systems, or cultural contexts, limiting the external validity of the models.

Cross-Sectional Assessment: Supportive care needs were measured at a single time point, whereas needs evolve throughout the cancer treatment trajectory; longitudinal models would better capture this dynamic.

Chinese Healthcare Context: Healthcare system structure, family support norms, and patient-provider communication patterns in China may differ substantially from Western settings, affecting model transferability.

Future Directions: Prospective validation across multiple institutions, extension to longitudinal need prediction, and integration of quality-of-life outcome data would strengthen the clinical utility of this approach.

TL;DR: Single-center, cross-sectional design limits generalizability; future work should validate models longitudinally across multiple institutions and cultural contexts.
Citation: Open Access, 2025. Available at: PMC12165907.