Artificial intelligence assistance significantly improves Gleason grading of prostate biopsies by pathologists

Mod Pathol 2021 Digital Pathology 6 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
The Problem of Gleason Grading Variability

The Gleason score is the most important tissue-based prognostic marker in prostate cancer. Assigned by a pathologist examining a biopsy under the microscope, it describes the architectural pattern of cancer glands and predicts how aggressively the tumor will behave. Scores are organized into five ISUP grade groups, with grade group 1 indicating the lowest risk and grade group 5 the highest.

Despite its clinical importance, Gleason grading suffers from significant inter-observer variability -- the same biopsy graded by different pathologists often receives different scores. Studies have shown substantial disagreement even among experienced pathologists, and the problem is worse at institutions without subspecialty-trained uropathologists. This variability directly affects treatment decisions: a patient graded as GG1 may be offered active surveillance while a GG3 patient receives immediate treatment.

Standalone deep learning AI systems have been developed that achieve pathologist-level performance on Gleason grading in benchmark evaluations. However, AI systems have known failure modes: they can be misled by artifacts such as ink on slides, non-prostate tissue, fixation artifacts, scanning defects, and rare cancer subtypes. Pathologists can recognize these problems immediately, but the AI cannot reliably flag them.

This gap motivated the central research question of this study: rather than asking whether AI can replace the pathologist, what happens when a pathologist uses AI as a real-time decision support tool? The combination of human visual expertise and AI computational analysis could produce a human-AI synergy that outperforms either working alone.

TL;DR: Gleason grading of prostate biopsies suffers from significant pathologist variability, and while AI systems can match average pathologist performance, whether AI assistance actively improves human grading had not previously been studied.
Pages 2-5
Study Design: Pathologists Grading With and Without AI

The study used a deep learning system previously developed on 5,759 H&E-stained biopsy slides from 1,243 patients, which had been independently validated against a panel of pathologists. A test set of 160 biopsies was assembled for the observer experiment, covering a carefully balanced distribution: 30 benign cases (including pitfalls such as partial atrophy, reactive atypia, and high-grade PIN), and equal numbers of grade groups 1 through 5 among the malignant cases.

A reference standard was established by three uropathologists with more than 20 years of subspecialty experience, who graded the biopsies individually and then resolved disagreements through a structured consensus protocol. Their agreement in the first round was very high (quadratically weighted Cohen's kappa of 0.925), providing a reliable gold standard against which pathologist performance could be measured.

An observer panel of 14 pathologists (11 board-certified and 3 pathology residents) from 12 independent laboratories in 8 countries first graded all 160 cases without AI assistance through an online digital viewer. At least 3 months later, they re-graded the same 160 cases alongside AI-generated overlays. The AI highlighted individual malignant glands using color coding: Gleason pattern 3 in yellow, pattern 4 in orange, and pattern 5 in red, along with a numerical prediction of the grade group and estimated volume percentages for each pattern.

A set of 60 previously unseen control cases was interspersed within the 160 cases to measure whether any improvement in the second read came simply from pathologists being more practiced rather than from the AI assistance itself. After completing the second read, all pathologists completed a questionnaire about how they experienced the AI feedback and which components they found most useful.

TL;DR: Fourteen pathologists from 8 countries graded 160 digitized prostate biopsies twice -- first without AI, then with color-coded Gleason pattern overlays and numerical predictions -- with a 3-month washout period between reads.
Pages 7-9
AI Assistance Significantly Improved Grading Accuracy

In the unassisted read, the median panel agreement with the expert reference standard was a quadratically weighted Cohen's kappa of 0.799. With AI assistance, this increased to 0.872 -- a 9.1% relative improvement that was statistically significant (Wilcoxon signed-rank test, p = 0.019). The standalone AI system itself achieved a kappa of 0.854 on the same dataset, meaning that the AI-assisted panel outperformed both the unassisted panel and the standalone AI system as a group.

At the individual level, 9 of 14 (64%) panel members improved with AI assistance, while 5 (36%) scored slightly lower -- but none decreased by more than 0.013 kappa points. Notably, 4 of the 5 who scored lower had already outperformed the AI system in the unassisted read, suggesting the AI assistance had limited additional benefit for the highest-performing pathologists. Of the 10 pathologists who scored below the AI in the first read, 9 (90%) improved in the assisted read to exceed the AI's performance.

Variability between panel members decreased with AI assistance. The interquartile range of individual kappa values dropped from 0.113 to 0.073, and the median pairwise agreement among panel members increased from a kappa of 0.737 to 0.859. Agreement on estimated tumor volume also became more consistent, with the pairwise correlation increasing from Pearson r = 0.744 to 0.780 and its interquartile range decreasing from 0.164 to 0.105. This standardizing effect is directly clinically relevant because it means that a patient's prognosis depends less on which pathologist happens to read their biopsy.

The largest improvements were seen for less experienced pathologists. Pathologists with fewer than 15 years of experience -- who tended to score below the AI system in the unassisted read -- benefited most substantially from the AI overlay. Pathologists with extensive experience, who already graded at or above the AI system's level, showed more modest or no improvement, suggesting that AI assistance functions primarily as an expertise equalizer.

TL;DR: AI-assisted pathologists achieved a group kappa of 0.872 versus 0.799 without AI, outperforming both the unassisted panel and the standalone AI system, with the greatest benefit seen for less experienced pathologists.
Page 9
External Validation Confirms the Benefit

To confirm the internal experiment results were not specific to one dataset, the panel was asked 6 months later to grade 87 biopsies from the Imagebase dataset -- a publicly available collection of prostate needle biopsies graded by 24 international experts specifically selected to represent a wide range of challenging tissue patterns.

In the unassisted read on these external cases, the panel's median pairwise agreement with the international expert panel was a kappa of 0.733. With AI assistance, this increased significantly to 0.786 (p = 0.003). The proportion of panel members outperforming the standalone AI system rose from 25% (3 out of 12) in the unassisted read to 83% (10 out of 12) in the assisted read.

Both the AI system and the panel scored somewhat lower on the external Imagebase dataset than on the internal test set. This was expected: the Imagebase cases were curated to be diagnostically difficult, the reference standard was based on independent microphotographs rather than whole-slide images, and the biopsies were scanned on different equipment than the AI was trained on. Despite this domain shift, AI assistance still produced a statistically significant improvement.

The consistency of the improvement across two independent datasets -- one from a single institution in the Netherlands and one from an international expert panel -- supports the conclusion that the observed benefit of AI assistance reflects a genuine and reproducible effect rather than an artifact of the specific test set or reference standard used.

TL;DR: On an independent external dataset graded by 24 international experts, AI assistance again significantly improved panel performance and the proportion of pathologists outperforming the standalone AI rose from 25% to 83%.
Pages 6, 7, 10
What Pathologists Found Most Useful

Questionnaire responses from the panel revealed that 79% of pathologists actively used the AI feedback during the second read. Despite the majority (57%) initially predicting that AI would not improve their performance, the results showed otherwise -- highlighting that pathologists may underestimate their potential to benefit from decision support tools.

Of all the AI output components, the gland-level Gleason pattern overlay was rated most useful. This visual overlay color-coded each individual malignant gland by its predicted pattern -- yellow for pattern 3, orange for pattern 4, red for pattern 5 -- providing spatially specific feedback that pathologists could immediately relate to the tissue they were examining. The AI's predicted grade group at the case level was rated least useful, likely because it was redundant with the numerical outputs already provided.

Pathologists reported that AI assistance did not distract them from the grading process. Instead, the majority indicated that grading became faster with AI assistance, even though time per case was not formally measured in this study. This subjective report aligns with findings from other pathology AI studies where AI-generated hotspot maps and detection overlays have been shown to reduce time-to-diagnosis in lymph node evaluation and mitosis counting tasks.

The fact that pathologists were not formally trained on the AI system before the experiment and still showed significant improvement suggests that the interface was sufficiently intuitive. The authors note that formal onboarding -- teaching pathologists about the AI system's known limitations and calibration -- could further increase the benefit, particularly by helping pathologists decide when to override versus follow the AI's suggestions.

TL;DR: The Gleason pattern color overlay was the most valued AI feature, and pathologists reported grading became faster with assistance -- suggesting AI support improves both accuracy and efficiency simultaneously.
Pages 10-11
Limitations and Clinical Implications

The study has several important limitations. The experiment used a case-level design where each biopsy was reviewed individually, whereas in clinical practice pathologists evaluate all biopsies from a patient together. Multi-biopsy patient-level assessment introduces the possibility of AI-assisted prioritization -- for example, automatically flagging which biopsy slides contain the highest-grade tumor -- which could further reduce workload and improve efficiency.

The 3-month washout period between reads was designed to minimize memory effects, and the inclusion of unseen control cases provided a check on practice effects. Statistical analysis excluding pathologists who attributed improvement to greater digital viewing experience still showed significant AI benefit. The external validation 6 months later, on a completely different dataset, further rules out experience as the primary explanation for improvement.

The study was conducted in a research setting with a carefully selected and balanced test set. Clinical validation requires demonstration that AI assistance improves grading on representative distributions of cases as encountered in daily practice, including patients with comorbidities, variable tissue quality, and unusual biopsy procedures. Multi-center prospective validation will be essential before clinical deployment.

Despite these limitations, the clinical implications are substantial. In regions with limited access to subspecialty uropathologists -- including many parts of sub-Saharan Africa, South America, and rural areas globally -- an AI-assisted grading system could elevate the quality of prostate cancer diagnosis toward specialist-level consistency. Reduced observer variability means that treatment decisions will depend less on the luck of which pathologist reads the biopsy, making the Gleason score a more reliable prognostic marker for individual patients.

TL;DR: While multi-center prospective validation is still needed, AI-assisted Gleason grading has immediate potential to standardize prostate cancer diagnosis across institutions and enable specialist-level accuracy in settings without subspecialty pathology expertise.
Citation: Open Access, . Available at: PMC7897578.