GLENDA: Gynecologic Laparoscopy Endometriosis Dataset
- GLENDA is a disease-specific dataset with region-based annotations extracted from over 400 surgical videos for endometriosis lesion localization.
- The dataset comprises more than 25,000 images, including both pathological and non-pathological frames, enabling robust lesion classification and tracking.
- It focuses on four key lesion locations—peritoneum, ovary, uterus, and DIE—supporting tasks from binary classification to detailed segmentation.
GLENDA, the Gynecologic Laparoscopy ENdometriosis DAtaset, is a publicly released image dataset for endometriosis in gynecologic laparoscopy. It was introduced as the first publicly released image dataset specifically focused on endometriosis in gynecologic laparoscopy, and it contains region-based annotations of endometriosis lesions together with non-pathological material, preserving links between distributed images and their source laparoscopic video segments. In practice, GLENDA functions as a curated resource for lesion detection and localization, pathology classification, and temporally informed analysis in minimally invasive gynecologic surgery (Leibetseder et al., 29 Aug 2025).
1. Clinical and technical rationale
GLENDA is situated at the intersection of gynecologic minimally invasive surgery and surgical video analysis. Gynecologic laparoscopy is performed via a live feed of a patient’s abdomen, and the resulting recordings are important for treatment planning and follow-up, case documentation, education and surgical training, and legal documentation. The dataset addresses the fact that manual analysis of such recordings is extremely time-consuming and that hospital video archives are large, heterogeneous, and difficult to search (Leibetseder et al., 29 Aug 2025).
The clinical target is endometriosis, described in the dataset paper as a benign but often very painful condition characterized by the presence of uterine-like tissue outside the uterus. Because diagnosis and treatment are typically performed laparoscopically, the operative video stream provides a direct visual substrate for computer vision methods. GLENDA was therefore designed to support lesion detection and localization, classification, and temporal analysis, rather than only procedural workflow analysis or instrument recognition (Leibetseder et al., 29 Aug 2025).
A notable aspect of the resource is its disease specificity. In the broader ecosystem of laparoscopic and endoscopic datasets, many resources emphasize cholecystectomy workflow, general content analysis, or gastrointestinal endoscopy. GLENDA instead targets endometriosis lesion localization in gynecologic laparoscopy, making it pathology-centric rather than workflow-centric (Leibetseder et al., 29 Aug 2025).
2. Data composition and curation strategy
GLENDA was assembled from a source pool of 400+ full gynecologic laparoscopy surgery videos. From this archive, the released dataset selects 300+ video segments and individual frames, yielding over 25,000 images, including 12,000+ “positive” / pathological images with visible endometriosis and 13,000+ “negative” / no visible endometriosis images (Leibetseder et al., 29 Aug 2025).
Although distributed as an image-based dataset, GLENDA preserves video context. It contains both annotated single frames and short sequences of consecutive frames extracted from the original videos. In pathology sequences, one keyframe is annotated with lesion regions; in “no pathology” sequences, no regions are annotated. The sequences were explicitly designed so that annotated pathology regions remain visible across the sequence with minimal camera motion, which makes them suitable for temporal tracking and annotation propagation (Leibetseder et al., 29 Aug 2025).
The curation strategy is anatomically focused. The first release concentrates on four lesion locations relevant to standard endometriosis classification concepts:
- Peritoneum
- Ovary
- Uterus
- DIE (Deep Infiltrating Endometriosis)
It also includes a fifth no pathology category consisting of sequences with no visible endometriosis. The paper states that the selection was shaped by expert availability, the rarity of certain lesion locations and severity combinations, and a practical decision to focus on a subset of the more than 50 possible location-by-severity combinations implied by rASRM and Enzian (Leibetseder et al., 29 Aug 2025).
The paper does not provide an explicit patient count, exact camera models, frame rates, or acquisition resolutions for GLENDA. It does state that the images are extracted video frames saved as standard image files and that no explicit pre-processing such as color normalization or cropping is specified beyond extraction and structuring (Leibetseder et al., 29 Aug 2025).
3. Annotation model, taxonomy, and distributed structure
GLENDA uses region-based annotations rather than only frame-level labels. Annotations were created with the Endoscopic Concept Annotation Tool (ECAT), using closed free-hand drawings, polygons, and rectangles. Each annotated region encloses a contiguous lesion area and is assigned to one of the four pathological categories. In the distributed dataset, these regions are provided as binary mask images with black background and white lesion pixels; importantly, the masks are stored per annotation, not as pre-merged per-frame multi-object masks (Leibetseder et al., 29 Aug 2025).
The label taxonomy is intentionally flat. The pathology classes are peritoneum, ovary, uterus, and die, together with the non-pathology class no_pathology. The dataset is informed by the rASRM and Enzian clinical schemes, but GLENDA’s first release does not encode full rASRM or Enzian severity scores. Its labels are anatomical-location labels rather than complete surgical staging descriptors (Leibetseder et al., 29 Aug 2025).
Annotation density is nontrivial. A single frame can contain up to 9 annotations, and as many as 3 different pathology classes per frame. For sequences, one or more keyframes are annotated while the corresponding lesion remains visible across neighboring frames, which supports tracking-style use cases and sequence-level model design (Leibetseder et al., 29 Aug 2025).
The distributed archive includes a DS directory, a Readme.md, and statistics/annotations.csv and statistics/images.csv. Frames are organized under pathology and no-pathology branches, while lesion masks are stored under sequence- and frame-specific paths using filenames of the form CLASS_a_AID.png. This organization allows reconstruction of per-frame masks, polygons, or bounding boxes from the individual binary annotation files. GLENDA is publicly available at http://www.itec.aau.at/ftp/datasets/GLENDA (Leibetseder et al., 29 Aug 2025).
4. Statistical profile and visual characteristics
The released dataset contains 138 sequences and 25,682 frames in total, of which 302 frames carry at least one annotation. The pathological distribution is strongly skewed toward peritoneal disease, while uterus is the sparsest class. The table below reproduces the main per-category counts reported for GLENDA (Leibetseder et al., 29 Aug 2025).
| Category | Annotations | Frames |
|---|---|---|
| Peritoneum | 402 | 6470 |
| Ovary | 51 | 2478 |
| Uterus | 14 | 475 |
| DIE | 53 | 2821 |
| No pathology | 0 | 13,438 |
| Total | 520 | 25,682 |
The same summary also reports 203 annotated frames for peritoneum, 48 for ovary, 8 for uterus, 43 for DIE, and a dataset-level maximum of 9 annotations per frame, with up to 3 pathology classes per frame. The high-level pathology versus non-pathology split is therefore comparatively balanced, but the pathological subclasses are not: peritoneum dominates, uterus is extremely small, and no-pathology sequences are numerous and long (Leibetseder et al., 29 Aug 2025).
GLENDA’s visual regime is explicitly realistic rather than normalized. The paper emphasizes a laparoscopic color spectrum dominated by red, yellow, and white, with comparatively little green or blue. Peritoneum appears as a large membrane covering much of the pelvic cavity; ovaries are often oval-shaped and white; the uterus is pear-shaped and appears in various red shades; DIE is described as visually heterogeneous and often similar to other anatomical structures. The distributed material retains instruments, varying camera angles and distances, smoke, blood and fluid, occlusions, partial organ views, lighting changes, and specular reflections (Leibetseder et al., 29 Aug 2025).
These characteristics are directly relevant to model design. They imply that GLENDA is not a sanitized lesion dataset but a realistic operative corpus in which pathology recognition must be robust to intraoperative confounders and anatomical variation. This also explains why the paper treats sequence-level splitting and class balancing as critical methodological concerns (Leibetseder et al., 29 Aug 2025).
5. Supported tasks and downstream reuse
The dataset paper identifies several intended use cases. At the simplest level, GLENDA supports binary classification of pathology versus no pathology by contrasting DS/pathology/frames with DS/no_pathology/frames. It also supports multi-class classification of lesion location, localization and segmentation via the distributed region masks, temporal analysis and tracking on the short sequences, weakly supervised propagation from annotated keyframes to neighboring frames, and surgical video indexing and decision support based on lesion-aware retrieval from archived recordings (Leibetseder et al., 29 Aug 2025).
A direct downstream example is the demo paper “Post-surgical Endometriosis Segmentation in Laparoscopic Videos”, which states that it “custom-create[s] a single-class lesion dataset from refining parts of the more extensive and multi-class Gynecologic Laparoscopy Endometriosis Dataset (GLENDA)”. That work derives a focused subset for dark endometrial implants, comprising 160 frames taken from more than 100 patient cases exhibiting endometriosis with over 350 region-based endometrial implant annotations, and uses it to train a Mask R-CNN with ResNet-101, 50 epochs, SGD, and learning rate 0.001. The resulting system performs frame-by-frame lesion segmentation and produces annotated laparoscopic videos with multi-colored overlays and a temporal indication bar; its best model, trained with rotation and cropping, achieved 0.642 mAP@0.50IoU and 0.324 mAP averaged over IoU thresholds from 0.50 to 0.95 (Leibetseder et al., 14 Oct 2025).
This downstream refinement is important because it shows that GLENDA is not only a static archival release. It can be subsetted into task-specific lesion corpora, used for practical segmentation systems, and connected to video browsing workflows via annotated video outputs and optional JSON metadata (Leibetseder et al., 14 Oct 2025).
6. Position among related datasets, constraints, and projected extensions
Within the broader gynecologic laparoscopy dataset landscape, GLENDA occupies a specialized role. In the comparison reported by GynSurg, GLENDA is characterized as a public dataset for “Detection and Localization” with 25,682 frames and 5 pathological categories. GynSurg, by contrast, is described as a multi-task gynecologic laparoscopy dataset centered on action recognition, side-effect detection, instrument segmentation, and anatomy segmentation rather than direct lesion labeling. In that sense, GLENDA is pathology-centric, whereas GynSurg is workflow- and scene-centric; the two are complementary rather than interchangeable (Nasirihaghighi et al., 12 Jun 2025).
Several limitations are explicit in GLENDA’s own description. First, there is pronounced class imbalance, especially the dominance of peritoneum and the very small uterus class. Second, the dataset contains strong temporal correlation because many frames come from short consecutive sequences; the paper warns that train/test splits made at frame level can cause severe data leakage and even trivial validation behavior. Third, coverage is intentionally incomplete: only four anatomical lesion classes are included, without full rASRM or Enzian severity modeling. Fourth, the material comes from a single institutional environment, which raises transferability questions across hospitals and devices. Fifth, annotations were produced by a small set of experts, but no inter-annotator agreement statistics are reported. Finally, GLENDA does not include explicit organ or instrument segmentation; only lesion regions are labeled (Leibetseder et al., 29 Aug 2025).
The planned extension path is equally explicit. The dataset paper states the intention to add other lesion locations, including tubes, ligaments, bowel/rectum, bladder, and ureter, with the long-term aim of covering the full rASRM and Enzian location sets. It also proposes adding severity annotations on a three-level scale and introducing “endometriosis suspicion” classes for uncertain or borderline cases, which would explicitly model clinical ambiguity. Continued annotation of more surgeries and possible expansion beyond a single center are also identified as future directions (Leibetseder et al., 29 Aug 2025).
Taken together, these properties define GLENDA as a foundational yet deliberately bounded dataset: a publicly accessible, disease-specific, region-annotated laparoscopic corpus that foregrounds endometriosis lesion localization, exposes the methodological difficulties of realistic surgical video, and has already served as the basis for more focused post-surgical segmentation systems (Leibetseder et al., 29 Aug 2025).