A Large-Scale Whole-Slide Image Dataset for Detecting Malignancy in Endometrial Biopsies

Gigascience 2025 AI 6 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Why Endometrial Pathology Needs a Dedicated AI Dataset

Endometrial cancer is diagnosed by pathological examination of tissue from uterine biopsies. For artificial intelligence tools to reliably assist in this diagnostic process, they must be trained on large, representative collections of digitized tissue slides with expert-verified labels.

Most publicly available pathology datasets focus on common cancers such as breast or lung cancer. Endometrial pathology datasets are scarce, limiting the development and validation of AI tools for this common gynecologic malignancy. Researchers building AI models have been constrained by small, proprietary datasets that cannot be shared or reproduced.

This study addresses that gap by releasing a large, publicly available whole-slide image (WSI) dataset specifically designed for detecting malignancy in endometrial biopsy specimens, with rich expert annotations and deliberately varied acquisition conditions to maximize generalizability.

TL;DR: A large public dataset of annotated endometrial whole-slide images is needed to enable reproducible AI development for gynecologic pathology.
Pages 2-3
Dataset Composition: Scale and Diversity

The dataset contains 2,909 whole-slide images from the NHS Greater Glasgow and Clyde pathology service, one of the largest routine pathology services in the United Kingdom. This scale is substantially larger than most prior endometrial pathology datasets and comparable to the leading breast and lung pathology benchmarks.

Slides are classified into three diagnostic categories: malignant (860 slides), representing confirmed endometrial carcinoma; other/benign (1,867 slides), covering hyperplasia, polyps, atrophic endometrium, and other non-cancer diagnoses; and insufficient (182 slides), where the tissue sample was inadequate for diagnosis.

Including the "insufficient" category is a practical feature rarely seen in academic datasets. In real clinical practice, a significant proportion of biopsies yield inadequate material, and an AI system that cannot recognize insufficient specimens would be unsafe for clinical deployment.

TL;DR: The dataset contains 2,909 slides in three classes including an 'insufficient' category, reflecting the full range of real-world biopsy outcomes.
Pages 3-5
Expert Annotation Process and Quality Control

All slides were annotated by consultant histopathologists -- senior specialist physicians -- using the original clinical pathology reports as the reference standard. This approach ensures that labels reflect expert consensus diagnostic practice rather than research-setting annotations that may not align with clinical norms.

Slides were deliberately sourced from eight different laboratories within the NHS Greater Glasgow system. Each laboratory uses slightly different staining protocols, scanner types, and preparation methods, introducing the kind of technical variation that AI systems will encounter when deployed across multiple hospital sites.

This multi-laboratory design is a key strength. AI models trained on data from a single lab often fail when moved to a new institution due to domain shift -- systematic differences in how slides look across sites. Training on multi-site data from the outset produces more robust and generalizable models.

TL;DR: Expert pathologist annotations from 8 different labs ensure the dataset reflects both diagnostic accuracy and realistic technical variation across hospital sites.
Pages 5-6
Slide Acquisition and Digital Infrastructure

All slides were digitized at 40x magnification equivalent using whole-slide scanners, producing images with gigapixel-scale resolution. At this resolution, individual cell nuclei, glandular structures, and tissue architecture are clearly visible, providing the detail necessary for pathological assessment.

The dataset is released with standard metadata including patient age range, biopsy indication, and scanner model. This information allows researchers to investigate whether AI model performance varies with patient demographic factors or scanning equipment -- important for understanding potential biases.

WSIs are provided in standard SVS or TIFF formats compatible with common pathology analysis platforms, lowering the barrier to entry for research teams without specialist data engineering resources. Accompanying code for patch extraction and basic preprocessing is also provided.

TL;DR: High-resolution 40x digitization with standard file formats and accompanying preprocessing code makes the dataset accessible for research teams worldwide.
Pages 6-7
Enabling Reproducible AI Research in Endometrial Pathology

A large, publicly available benchmark enables fair comparison between AI methods. When different research groups train and test their models on the same dataset using the same splits, their results can be directly compared, accelerating progress and reducing wasted effort from duplicate development of similar approaches.

The dataset is large enough to support both supervised learning approaches (which require many labeled examples) and evaluation of self-supervised or semi-supervised methods that aim to reduce annotation requirements. This flexibility makes it useful across a wide range of AI research paradigms.

Practically, AI tools trained on this dataset could support pathologists by triaging cases, flagging urgent malignant cases for priority review, and filtering out insufficient specimens early in the workflow -- directly reducing diagnostic bottlenecks in busy gynecologic pathology services.

TL;DR: This open dataset enables reproducible benchmarking of AI pathology methods and provides a foundation for tools that could triage endometrial biopsies in clinical practice.
Pages 7-8
Filling a Critical Gap in Pathology AI Resources

The release of this 2,909-slide dataset fills a long-standing gap in publicly available resources for endometrial pathology AI research. Its scale, multi-site origin, expert annotation, and inclusion of all clinically relevant diagnostic categories make it one of the most comprehensive gynecologic pathology datasets available.

Future expansions could include graded annotations of tissue regions, molecular subtype information from genomic profiling, and linkage to clinical outcomes such as recurrence and survival -- transforming the dataset from a diagnostic benchmark into a resource for prognostic AI development.

By making this dataset openly available, the authors are contributing to a foundation for AI-assisted endometrial cancer diagnosis that is reproducible, externally validated, and ultimately translatable to patient benefit in pathology services globally.

TL;DR: This open dataset establishes a new benchmark for endometrial pathology AI, with the scale, diversity, and clinical realism needed to produce deployable diagnostic tools.
Citation: Open Access, 2025. Available at: PMC12751089.