PD-L1 Scoring Models for Non-Small Cell Lung Cancer in China: Current Status, AI-Assisted Solutions and Future Perspectives

Thorac Cancer 2025 AI 5 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
The PD-L1 Testing Crisis in China's NSCLC Practice

Why PD-L1 Testing Matters PD-L1 expression on tumor cells is the primary biomarker used to determine whether NSCLC patients qualify for immune checkpoint inhibitor (ICI) therapy. The tumor proportion score (TPS) - the percentage of tumor cells with positive PD-L1 staining - determines drug eligibility and expected treatment benefit.

China's Unique Challenge China has approved 11 NMPA-approved anti-PD-1/PD-L1 monoclonal antibodies for NSCLC, each requiring PD-L1 testing with different companion diagnostic kits. This creates a complex landscape where the same tumor may yield different PD-L1 scores depending on which kit, platform, and pathologist evaluates it.

Inter-Pathologist Variability Problem Manual PD-L1 scoring by pathologists is known to be highly variable. Studies show significant discordance in TPS categorization (0%, 1-49%, 50%+) between pathologists - the very thresholds that determine treatment decisions. In China, the problem is amplified by the number of drugs requiring separate testing.

Review Scope and Purpose This comprehensive review article examines the current landscape of PD-L1 testing in China, catalogs existing AI-assisted scoring solutions, evaluates their performance, and proposes a roadmap for standardizing AI-based PD-L1 evaluation in Chinese clinical practice.

TL;DR: This review examines the complex PD-L1 testing landscape in China with 11 approved ICI drugs, analyzing current challenges in scoring variability and cataloging AI-assisted solutions for standardization.
Pages 2-3
PD-L1 Scoring Challenges Across Multiple Antibody Kits

Multiple Companion Diagnostics Each approved anti-PD-L1 drug in China uses a different antibody clone and platform for PD-L1 testing. The 22C3, 28-8, SP142, and SP263 clones all detect PD-L1 but with different staining intensities and cellular compartment specificity. A patient whose TPS is 48% on one assay may score 52% on another - changing their treatment eligibility.

Combined Positive Score (CPS) vs. TPS Beyond TPS (counting tumor cells only), some drugs require the Combined Positive Score (CPS), which includes immune cells and tumor cells. Accurately scoring CPS requires distinguishing tumor cells from infiltrating immune cells - a distinction that is particularly challenging in manually stained sections and requires expertise.

Laboratory Infrastructure Gaps While top-tier hospitals in major Chinese cities have advanced pathology departments, many regional hospitals lack standardized equipment, trained pathologists, and quality control programs. This creates inequality in PD-L1 testing quality across China's vast geographic landscape.

Current AI Solutions Reviewed The review catalogs AI-based PD-L1 scoring systems already in use or under evaluation in China, including cell-level detection algorithms, WSI analysis platforms, and hybrid human-AI scoring workflows. Performance across systems is compared using ICC and concordance metrics.

TL;DR: Multiple PD-L1 antibody clones with different scoring thresholds, combined with TPS vs. CPS distinctions and variable lab infrastructure, create major standardization challenges that AI solutions aim to address.
Pages 4-6
AI-Assisted PD-L1 Scoring: Current Approaches

Cell Detection and Counting The most straightforward AI approach uses deep learning models to detect and count individual PD-L1 positive and negative tumor cells in stained whole slide images (WSIs). These systems automate the manual counting process and reduce variability by applying consistent criteria across all cells in the slide.

Multiple Instance Learning (MIL) More advanced approaches use weakly supervised MIL methods that learn from slide-level TPS labels without requiring cell-by-cell annotation. MIL processes WSIs by dividing them into patches and learning which regions are most informative for TPS prediction - enabling training at scale without expensive annotation efforts.

Concordance with Pathologists Reviewed AI systems generally achieve intraclass correlation coefficients (ICC) of 0.85 to 0.95 when compared to expert pathologist TPS scores. This is comparable to or better than inter-pathologist concordance, suggesting AI can serve as a standardizing tool rather than just an efficiency tool.

Cross-Platform Harmonization Some AI systems have been trained across multiple PD-L1 antibody platforms and can produce harmonized scores - enabling comparison of PD-L1 status detected with different companion diagnostic kits. This could eventually allow a single AI-scored TPS to guide treatment decisions across multiple drugs.

TL;DR: AI PD-L1 scoring systems using cell detection algorithms or MIL achieve ICC 0.85-0.95 with pathologists and show potential for cross-platform harmonization, reducing inter-pathologist variability.
Pages 7-8
Standardizing PD-L1 Testing Through AI Adoption

Addressing Borderline Cases The highest clinical impact of AI scoring is at borderline TPS thresholds (1% and 50%), where reclassification has direct treatment consequences. AI systems that provide continuous TPS estimates with confidence intervals could flag borderline cases for expert review, reducing misclassification at critical cutoffs.

Quality Assurance and Auditing AI scoring systems can serve as quality assurance tools - flagging slides that deviate significantly from AI scores for pathologist re-review. This creates a safety net in routine practice and enables ongoing quality monitoring of pathologist performance across institutions.

Scalability for Screening Programs As NSCLC screening programs expand in China, the volume of PD-L1 tests requiring expert review will grow rapidly. AI-assisted pre-screening - automatically scoring straightforward cases and routing complex ones to experts - is essential for sustainable throughput.

Regulatory and Reimbursement Pathways AI PD-L1 scoring tools require NMPA approval as medical devices and integration into national reimbursement frameworks. The review discusses the regulatory pathway requirements and the evidence needed to support AI tool approval in the Chinese medical device regulatory system.

TL;DR: AI PD-L1 scoring could standardize borderline case classification, serve as institutional quality assurance, enable scalable screening programs, and requires NMPA device approval for clinical deployment.
Pages 9-10
Building a Standardized AI PD-L1 Ecosystem in China

National Standards Development The review calls for national PD-L1 testing standards that define minimum performance requirements for AI-assisted scoring tools, establish reference datasets for validation, and create certification processes for laboratories using AI systems. Pan-China standardization would reduce regional disparities in testing quality.

Multi-Antibody Training Datasets Future AI systems should be trained on large, diverse datasets covering all 11 approved PD-L1 antibody clones in China. Multi-clone training would enable a single AI model to produce harmonized, platform-agnostic TPS scores - simplifying clinical decision-making.

Integration with Treatment Outcomes Beyond concordance with pathologist scores, future AI tools should be validated against actual treatment outcomes - whether patients with AI-scored TPS achieve the expected survival benefit from ICIs. Outcome-linked validation would establish clinical utility rather than just analytical concordance.

Multimodal Biomarker Integration Future platforms should integrate PD-L1 TPS with other biomarkers from the same WSI - tumor mutational burden estimation from histology, tumor microenvironment immune cell quantification, and spatial biology analysis. A single AI analysis of the pathology slide could yield a comprehensive immunotherapy eligibility report.

TL;DR: National standardization, multi-antibody AI training, outcome-linked validation, and multimodal biomarker integration are the key priorities for building a robust AI PD-L1 testing ecosystem in China.
Citation: Open Access, 2025. Available at: PMC11973252.