---
title: Retinal Disease Dataset Overview
url: https://www.emergentmind.com/topics/retinal-disease-dataset
type: topic
---

# Retinal Disease Dataset Overview

Searching arXiv for the cited retinal dataset papers to ground the article.
Retinal disease datasets are curated collections of ophthalmic images or videos assembled for classification, segmentation, localization, fairness analysis, explainability, and multimodal transfer. In current research, the term covers color fundus photographs, optical coherence tomography (OCT) B-scans, OCT angiography (OCTA), infrared reflectance images, handheld fundus video, and ocular ultrasound clips. The field now includes small expert-curated datasets, integrated cross-source corpora, synthetic datasets, and multimodal benchmarks that explicitly separate training-time and test-time modality requirements [2602.19324] [2508.15986] [2312.08255] [2412.09402].

## 1. Modalities, scope, and representative resources

Retinal disease datasets differ first by imaging modality and only then by task definition. Color fundus datasets dominate screening and multi-disease classification; OCT datasets emphasize retinal layer structure and macular pathology; OCTA datasets target microvascular changes; infrared reflectance datasets support artery–vein analysis; handheld video preserves temporal vascular dynamics; and ocular ultrasound datasets address conditions such as retinal detachment when optical media or access constraints limit fundus-based examination [2405.14137] [2409.04137] [2203.14928] [2307.06577] [2508.04735].

| Dataset | Modality | Scope |
|---|---|---|
| Retinal OCT Image Classification – C8 [2602.19324] | OCT | 24,000 labeled OCT images across eight conditions |
| SynFundus-1M [2508.15986] | Color fundus | over 1,000,000 synthetic fundus images; 11 multi-label diseases |
| OCTDL [2312.08255] | OCT B-scan | 2064 images from 821 patients |
| MultiEYE [2412.09402] | Fundus + OCT | 58,036 fundus images and 45,923 OCT B-scans; nine classes |
| RFMiD [2512.04967] | Color fundus | over 45,000 images spanning 46 retinal disease categories |
| RVD [2307.06577] | Fundus video | 635 smartphone-based fundus videos with spatial and temporal annotations |
| Harvard-GF [2306.09264] | RNFLT + 3D OCT | 3,300 subjects with balanced racial groups for glaucoma |
| ERDES [2508.04735] | Ocular ultrasound video | 5,381 deidentified clips for RD and macular status |

This diversity has produced several distinct dataset genres. Some datasets are disease-specific, such as HYAMD for age related macular degeneration and PALM for pathologic myopia [2505.04230] [2305.07816]. Others are broad diagnostic collections, such as EyeNet with 1,747 retinal images across 52 disease classes and MuReD with 2,208 fundus images and 20 multi-label classes [1808.05754] [2207.02335]. A further category consists of harmonization and benchmarking resources, such as IRFundusSet and MultiEYE, which reorganize heterogeneous sources into a unified experimental framework [2402.11488] [2412.09402].

## 2. Fundus photograph datasets and disease-label design

Color fundus datasets span binary screening, single-label multi-class diagnosis, and multi-label pathology detection. Disease-specific datasets include HYAMD, a longitudinal collection of 1,560 Digital Fundus Images from 325 patients examined between 2021 and 2024 for AMD identification and staging, and PALM, which provides 1,200 images from 720 subjects with pathologic myopia labels plus optic disc masks, fovea localization, and lesion masks for patchy retinal atrophy and retinal detachment [2505.04230] [2305.07816]. HYAMD is clinically notable because AMD labels were assigned following a full clinical assessment supported by OCT and OCT angiography; however, the manuscript contains discrepant cohort totals, with the abstract stating 325 patients while the methods section totals 380, and advises reliance on the labels file for exact counts [2505.04230].

Broader diagnostic fundus datasets make different choices about label space. EyeNet was curated from the Retina Image Bank and contains 1,747 retinal images across 52 disease classes with a 70% training, 10% validation, and 20% test split, emphasizing multi-class diagnosis under sparse data [1808.05754]. MuReD was constructed from ARIA, STARE, and the RFMiD training set, filtered to a minimum of 30 samples per retained label, and finalized at 20 classes, including DR, NORMAL, ARMD, BRVO, CSR, HTR, and OTHER [2207.02335]. RFMiD itself is multi-label, contains over 45,000 images spanning 46 retinal disease categories, and includes metadata such as patient ID, eye laterality, and image quality [2512.04967]. This difference is consequential: single-label datasets collapse coexisting disease into one target, whereas multi-label datasets explicitly model co-occurrence.

Synthetic and integrated fundus datasets address scale and fragmentation in different ways. SynFundus-1M contains over 1,000,000 synthetic fundus images generated with a Denoising Diffusion Probabilistic Model trained on 1.3 million private, authentic clinical fundus images, and supports eleven-disease multi-label classification [2508.15986]. IRFundusSet instead consolidates ten public fundus sources totaling 46,064 images, curates 25,406 of them for a harmonized `is_normal` label, and identifies 3,515 as healthy after manual review [2402.11488]. A central implication is that “source non-pathological” and “healthy” are not equivalent categories. IRFundusSet makes this explicit by distinguishing `src_is_normal` from a manually curated `is_normal`, and by reporting that the healthy rate in the curated set is 14.0%, substantially below the source-averaged “non-pathological” rate of 33.1% [2402.11488].

Fundus datasets also underpin foundation-model evaluation. RET-CLIP was pretrained on RET-Clinical, a patient-level image–report corpus of 193,865 patients from Beijing Tongren Hospital, and then evaluated on eight public datasets spanning diabetic retinopathy, glaucoma, multiple disease diagnosis, and multi-label classification, including IDRID, APTOS, PAPILA, JSIEC, RFMiD, and ODIR [2405.14137]. This suggests that “retinal disease dataset” increasingly refers not only to standalone benchmarks but also to transfer environments for representation learning.

## 3. OCT, OCTA, and multimodal retinal datasets

OCT-centered datasets foreground retinal microstructure. The Retinal OCT Image Classification – C8 dataset contains 24,000 labeled OCT images across eight conditions—Diabetic Macular Edema, Choroidal Neovascularization, Drusen, Central Serous Retinopathy, Macular Hole, Diabetic Retinopathy, Age-related Macular Degeneration, and Normal—and was used in RetinaVision after resizing all images to 224×224 pixels [2602.19324]. In that study, Xception reached 95.25% test accuracy and InceptionV3 94.82%, while Grad-CAM, LIME, and occlusion sensitivity were used to visualize discriminative retinal structures [2602.19324]. The same paper notes that the class distribution is “sufficiently representative,” but exact per-class counts and split ratios are not reported.

OCTDL represents a different design philosophy: smaller scale, denser metadata, and explicit patient-level leakage control. It contains 2064 high-resolution OCT B-scan images centered on the fovea from 821 patients, acquired with an Optovue Avanti RTVue XR using raster scanning protocols with dynamic scan length and image resolution [2312.08255]. The label space includes AMD, DME, ERM, RAO, RVO, VID, and Normal, and metadata include `file_name`, `disease`, `subcategory`, `condition`, `patient_id`, `eye`, `sex`, `year`, `image_width`, and `image_height` [2312.08255]. The recommended split is 60:10:20 at the patient level, explicitly ensuring that all images of a patient reside in a single subset [2312.08255].

OCTA datasets target a narrower but clinically distinct signal. The OCTA dataset for diabetic retinopathy contains 268 retinal images from 179 individuals, acquired using a nonmydriatic Optovue Avanti Edition machine with 8 × 8 mm wide scans centered on the macula, and annotated into No DR, Mild NPDR, and Moderate NPDR [2409.04137]. Images were exported as JPEG files at 1596 × 990 pixels, and poor-quality, blurry, artifact-laden, post-laser, and proliferative DR cases were excluded [2409.04137]. This makes the dataset specifically suitable for early-stage DR analysis rather than full-spectrum DR staging.

MultiEYE formalizes a multimodal regime in which OCT is available during training but not required during inference. It combines 58,036 fundus images and 45,923 OCT B-scans across nine single-label disease categories—Normal, dAMD, CSC, DR, GLC, MEM, MYO, RVO, and wAMD—with patient-identity-level splits maintained at approximately 6:2:2 [2412.09402]. Fundus and OCT are explicitly unpaired; the benchmark supports OCT-enhanced disease recognition from fundus images through concept-guided distillation rather than paired multimodal fusion [2412.09402]. This suggests a shift from paired multimodal datasets toward clinically asymmetric training regimes, where expensive structural modalities act as teachers for more ubiquitous imaging.

## 4. Temporal, infrared, and ultrasound resources

Retinal disease datasets are no longer limited to static optical images. RVD is described as the first publicly available fundus video dataset for vessel segmentation using handheld, smartphone-based acquisition, comprising 635 RGB fundus videos from 415 participants aged 50–75 years across four clinics [2307.06577]. It provides three levels of spatial annotations—binary vessel masks, general artery–vein masks, and fine-grained artery–vein masks—and temporal annotations for spontaneous retinal venous pulsation, including presence, duration, and peak/trough frames [2307.06577]. Because handheld capture introduces eye motion, video jitter, motion blur, nonuniform illumination, and operator variability, the dataset exposes a domain shift relative to bench-top image datasets [2307.06577].

RAVIR extends retinal vascular analysis into infrared reflectance imaging. It contains 46 Heidelberg Spectralis IR images at 768×768 pixels, 12.5 microns per pixel, with pixel-wise semantic masks for background, vein, and artery classes [2203.14928]. The pathology spectrum includes retinal vein occlusion, hypertensive retinopathy, peripapillary atrophy, diabetic retinopathy, isolated vessel tortuosity, high myopia, and media opacities [2203.14928]. The dataset was created to support artery–vein segmentation and vessel width measurement, and the SegRAVIR pipeline reports artery Dice 0.8287 and vein Dice 0.8301, alongside artery and vein width estimates of \(98.29\pm 5.48\) µm and \(112.34\pm 10.02\) µm respectively [2203.14928].

ERDES moves outside fundus/OCT imaging entirely and addresses retinal detachment in ocular ultrasound. It contains 5,381 deidentified ocular ultrasound video clips from unique participants, collected in the Emergency Department at the University of Arizona, and labeled for Non-RD, PVD, and RD; RD clips are further categorized as Macula_Intact or Macula_Detached, with anatomical subclasses such as TD, ND, and Bilateral [2508.04735]. For the benchmark tasks, PVD is excluded from binary Non-RD vs RD detection. On the test set, 3D U-Net achieved sensitivity 0.950 and accuracy 0.991 for Non-RD vs RD, and sensitivity 0.899 and accuracy 0.882 for Macula-Intact vs Macula-Detached [2508.04735]. A plausible implication is that retinal disease datasets now encode not only pathology labels but also time-critical triage variables.

## 5. Annotation practice, split strategy, and benchmarking methodology

Annotation protocols vary from image-level diagnosis to pixel-level masks and demographic metadata. OCTDL uses a multi-stage consensus process involving 7 trained medical students, two experienced clinical specialists, and final confirmation by the head of clinic expert [2312.08255]. PALM uses seven ophthalmologists plus one senior ophthalmologist for optic disc masks, fovea points, and lesion masks, with majority voting and senior quality control [2305.07816]. The OCTA DR dataset was graded by two experts—a board-certified ophthalmologist and a retina surgeon—while ERDES used three ocular ultrasound clinical experts plus a fourth expert for quality control and verification [2409.04137] [2508.04735]. Several datasets, however, do not report inter-rater agreement statistics.

Split strategy is a recurrent methodological fault line. Patient-level partitioning is explicit in OCTDL, PALM, MultiEYE, and Harvard-GF, and is recommended to avoid leakage across repeated images, visits, or eyes [2312.08255] [2305.07816] [2412.09402] [2306.09264]. By contrast, EyeNet reports a random 70/10/20 split and MuReD reports image-level train/validation organization without patient-level split control [1808.05754] [2207.02335]. HYAMD is longitudinal and includes repeated visits and both eyes, so cross-patient and eye-aware splitting is recommended in its usage guidance [2505.04230]. This suggests that reported accuracy is inseparable from split design, especially when datasets include repeated measures or bilateral images.

Benchmark methodology also reflects dataset structure. SynFundus-1M uses 5-fold multi-label stratified cross-validation and constructs a 66-dimensional meta-dataset from out-of-fold predictions of six base models across 11 diseases for XGBoost stacking [2508.15986]. RFMiD, in a few-shot setting, is used with balanced 5-way, 5-shot episodes and 2 query samples per class, trained on 100 episodes and evaluated on 1,000 test episodes [2512.04967]. SegImgNet applies five-fold cross-validation to AIROGS and e-ROP after resizing images to 256 × 256 pixels, using ROSE oversampling and Weighted Cross-Entropy to handle the severe imbalance of e-ROP [2503.00267]. RetinaVision applies CutMix and MixUp at 224×224, with Adam, categorical cross-entropy, batch size 32, learning rate 0.0001, 50 epochs, and early stopping [2602.19324].

Evaluation metrics track task type. Classification studies report Accuracy, Precision, Recall, F1, AUC, macro-AUC, and Kappa; segmentation datasets report Dice, IoU, mIoU, mAcc, and mFscore; few-shot and multi-label studies emphasize macro-averaged metrics under imbalance; and fairness benchmarks add DPD, DEO, DEOdds, ES-Acc, and ES-AUC [2508.15986] [2312.08255] [2306.09264]. Explainability has become part of benchmark design in some OCT and multimodal studies: RetinaVision uses Grad-CAM and LIME, while MultiEYE uses concept bottlenecks and image–concept similarity for OCT-assisted conceptual distillation [2602.19324] [2412.09402].

## 6. Generalization, fairness, and recurrent limitations

Several recurring limitations define the current retinal dataset literature. Class imbalance is explicit in RFMiD, OCTDL, e-ROP, ERDES, and C8; MuReD attempts to mitigate this by removing labels with fewer than 30 samples and reassigning them to OTHER, while few-shot work on RFMiD uses balanced episodic sampling and minority-focused augmentation [2512.04967] [2207.02335] [2503.00267]. Exact per-class counts and split ratios are missing in some otherwise widely used resources, including C8 and parts of HYAMD [2602.19324] [2505.04230].

Domain shift is equally pervasive. RetinaVision notes noise and variability from different acquisition systems as an ongoing challenge [2602.19324]. OCTDL emphasizes device specificity, since its scans come from Optovue Avanti whereas many public datasets use Spectralis or Cirrus [2312.08255]. RVD reports substantial gaps between handheld video data and bench-top vessel datasets [2307.06577]. SynFundus-1M demonstrates that models trained exclusively on synthetic data can generalize to real datasets, but with lower external AUC on DR than on AIROGS glaucoma or RFMiD, and explicitly recommends fine-tuning on smaller, expert-labeled real datasets for deployment [2508.15986]. This suggests that scale alone does not remove modality, device, or site dependence.

Fairness and label semantics are now treated as dataset-level design problems rather than post hoc concerns. Harvard-GF was built as a dedicated fairness dataset with 3,300 subjects, perfectly balanced racial groups, dual-modality imaging, and demographic metadata including race, gender, ethnicity, language, and marital status [2306.09264]. IRFundusSet addresses a different bias source: inconsistency in what public datasets call “healthy,” and therefore introduces a harmonized `is_normal` label after physical review across ten public cohorts [2402.11488]. Together, these resources show that retinal disease datasets are increasingly evaluated ոչ only by size or accuracy, but also by whether their labels, splits, and metadata support equitable and clinically interpretable learning.

A final misconception concerns what counts as a “retinal disease dataset.” The literature now includes disease-specific collections, integrated catalogs, synthetic generators, patient-level image–report corpora, cross-modal teacher–student benchmarks, vessel-segmentation resources, and time-resolved video datasets [2405.14137] [2402.11488] [2412.09402]. A plausible implication is that the field has moved from dataset release as isolated curation toward dataset design as experimental infrastructure: label ontology, modality asymmetry, temporal signal, demographic balance, and leakage control are increasingly first-class variables in retinal AI research.

Source: https://www.emergentmind.com/topics/retinal-disease-dataset