Retinal Disease Dataset Overview
- Retinal disease datasets are curated collections of diverse ophthalmic images (color fundus, OCT, OCTA, infrared, video, ultrasound) designed for disease classification and segmentation.
- They employ rigorous annotation protocols, patient-level splitting, and synthetic augmentation to ensure methodological rigor and fairness in performance evaluation.
- These datasets facilitate multimodal transfer learning and representation analysis, with benchmarks showing high accuracy and balanced demographic metrics for clinical applications.
Searching arXiv for the cited retinal dataset papers to ground the article. Retinal disease datasets are curated collections of ophthalmic images or videos assembled for classification, segmentation, localization, fairness analysis, explainability, and multimodal transfer. In current research, the term covers color fundus photographs, optical coherence tomography (OCT) B-scans, OCT angiography (OCTA), infrared reflectance images, handheld fundus video, and ocular ultrasound clips. The field now includes small expert-curated datasets, integrated cross-source corpora, synthetic datasets, and multimodal benchmarks that explicitly separate training-time and test-time modality requirements (Noor et al., 22 Feb 2026, Cao-Xue et al., 21 Aug 2025, Kulyabin et al., 2023, Wang et al., 2024).
1. Modalities, scope, and representative resources
Retinal disease datasets differ first by imaging modality and only then by task definition. Color fundus datasets dominate screening and multi-disease classification; OCT datasets emphasize retinal layer structure and macular pathology; OCTA datasets target microvascular changes; infrared reflectance datasets support artery–vein analysis; handheld video preserves temporal vascular dynamics; and ocular ultrasound datasets address conditions such as retinal detachment when optical media or access constraints limit fundus-based examination (Du et al., 2024, Bidwai et al., 2024, Hatamizadeh et al., 2022, Khan et al., 2023, Navard et al., 5 Aug 2025).
| Dataset | Modality | Scope |
|---|---|---|
| Retinal OCT Image Classification – C8 (Noor et al., 22 Feb 2026) | OCT | 24,000 labeled OCT images across eight conditions |
| SynFundus-1M (Cao-Xue et al., 21 Aug 2025) | Color fundus | over 1,000,000 synthetic fundus images; 11 multi-label diseases |
| OCTDL (Kulyabin et al., 2023) | OCT B-scan | 2064 images from 821 patients |
| MultiEYE (Wang et al., 2024) | Fundus + OCT | 58,036 fundus images and 45,923 OCT B-scans; nine classes |
| RFMiD (Khale et al., 4 Dec 2025) | Color fundus | over 45,000 images spanning 46 retinal disease categories |
| RVD (Khan et al., 2023) | Fundus video | 635 smartphone-based fundus videos with spatial and temporal annotations |
| Harvard-GF (Luo et al., 2023) | RNFLT + 3D OCT | 3,300 subjects with balanced racial groups for glaucoma |
| ERDES (Navard et al., 5 Aug 2025) | Ocular ultrasound video | 5,381 deidentified clips for RD and macular status |
This diversity has produced several distinct dataset genres. Some datasets are disease-specific, such as HYAMD for age related macular degeneration and PALM for pathologic myopia (Meisel et al., 7 May 2025, Fang et al., 2023). Others are broad diagnostic collections, such as EyeNet with 1,747 retinal images across 52 disease classes and MuReD with 2,208 fundus images and 20 multi-label classes (Yang et al., 2018, Rodriguez et al., 2022). A further category consists of harmonization and benchmarking resources, such as IRFundusSet and MultiEYE, which reorganize heterogeneous sources into a unified experimental framework (Githinji et al., 2024, Wang et al., 2024).
2. Fundus photograph datasets and disease-label design
Color fundus datasets span binary screening, single-label multi-class diagnosis, and multi-label pathology detection. Disease-specific datasets include HYAMD, a longitudinal collection of 1,560 Digital Fundus Images from 325 patients examined between 2021 and 2024 for AMD identification and staging, and PALM, which provides 1,200 images from 720 subjects with pathologic myopia labels plus optic disc masks, fovea localization, and lesion masks for patchy retinal atrophy and retinal detachment (Meisel et al., 7 May 2025, Fang et al., 2023). HYAMD is clinically notable because AMD labels were assigned following a full clinical assessment supported by OCT and OCT angiography; however, the manuscript contains discrepant cohort totals, with the abstract stating 325 patients while the methods section totals 380, and advises reliance on the labels file for exact counts (Meisel et al., 7 May 2025).
Broader diagnostic fundus datasets make different choices about label space. EyeNet was curated from the Retina Image Bank and contains 1,747 retinal images across 52 disease classes with a 70% training, 10% validation, and 20% test split, emphasizing multi-class diagnosis under sparse data (Yang et al., 2018). MuReD was constructed from ARIA, STARE, and the RFMiD training set, filtered to a minimum of 30 samples per retained label, and finalized at 20 classes, including DR, NORMAL, ARMD, BRVO, CSR, HTR, and OTHER (Rodriguez et al., 2022). RFMiD itself is multi-label, contains over 45,000 images spanning 46 retinal disease categories, and includes metadata such as patient ID, eye laterality, and image quality (Khale et al., 4 Dec 2025). This difference is consequential: single-label datasets collapse coexisting disease into one target, whereas multi-label datasets explicitly model co-occurrence.
Synthetic and integrated fundus datasets address scale and fragmentation in different ways. SynFundus-1M contains over 1,000,000 synthetic fundus images generated with a Denoising Diffusion Probabilistic Model trained on 1.3 million private, authentic clinical fundus images, and supports eleven-disease multi-label classification (Cao-Xue et al., 21 Aug 2025). IRFundusSet instead consolidates ten public fundus sources totaling 46,064 images, curates 25,406 of them for a harmonized is_normal label, and identifies 3,515 as healthy after manual review (Githinji et al., 2024). A central implication is that “source non-pathological” and “healthy” are not equivalent categories. IRFundusSet makes this explicit by distinguishing src_is_normal from a manually curated is_normal, and by reporting that the healthy rate in the curated set is 14.0%, substantially below the source-averaged “non-pathological” rate of 33.1% (Githinji et al., 2024).
Fundus datasets also underpin foundation-model evaluation. RET-CLIP was pretrained on RET-Clinical, a patient-level image–report corpus of 193,865 patients from Beijing Tongren Hospital, and then evaluated on eight public datasets spanning diabetic retinopathy, glaucoma, multiple disease diagnosis, and multi-label classification, including IDRID, APTOS, PAPILA, JSIEC, RFMiD, and ODIR (Du et al., 2024). This suggests that “retinal disease dataset” increasingly refers not only to standalone benchmarks but also to transfer environments for representation learning.
3. OCT, OCTA, and multimodal retinal datasets
OCT-centered datasets foreground retinal microstructure. The Retinal OCT Image Classification – C8 dataset contains 24,000 labeled OCT images across eight conditions—Diabetic Macular Edema, Choroidal Neovascularization, Drusen, Central Serous Retinopathy, Macular Hole, Diabetic Retinopathy, Age-related Macular Degeneration, and Normal—and was used in RetinaVision after resizing all images to 224×224 pixels (Noor et al., 22 Feb 2026). In that study, Xception reached 95.25% test accuracy and InceptionV3 94.82%, while Grad-CAM, LIME, and occlusion sensitivity were used to visualize discriminative retinal structures (Noor et al., 22 Feb 2026). The same paper notes that the class distribution is “sufficiently representative,” but exact per-class counts and split ratios are not reported.
OCTDL represents a different design philosophy: smaller scale, denser metadata, and explicit patient-level leakage control. It contains 2064 high-resolution OCT B-scan images centered on the fovea from 821 patients, acquired with an Optovue Avanti RTVue XR using raster scanning protocols with dynamic scan length and image resolution (Kulyabin et al., 2023). The label space includes AMD, DME, ERM, RAO, RVO, VID, and Normal, and metadata include file_name, disease, subcategory, condition, patient_id, eye, sex, year, image_width, and image_height (Kulyabin et al., 2023). The recommended split is 60:10:20 at the patient level, explicitly ensuring that all images of a patient reside in a single subset (Kulyabin et al., 2023).
OCTA datasets target a narrower but clinically distinct signal. The OCTA dataset for diabetic retinopathy contains 268 retinal images from 179 individuals, acquired using a nonmydriatic Optovue Avanti Edition machine with 8 × 8 mm wide scans centered on the macula, and annotated into No DR, Mild NPDR, and Moderate NPDR (Bidwai et al., 2024). Images were exported as JPEG files at 1596 × 990 pixels, and poor-quality, blurry, artifact-laden, post-laser, and proliferative DR cases were excluded (Bidwai et al., 2024). This makes the dataset specifically suitable for early-stage DR analysis rather than full-spectrum DR staging.
MultiEYE formalizes a multimodal regime in which OCT is available during training but not required during inference. It combines 58,036 fundus images and 45,923 OCT B-scans across nine single-label disease categories—Normal, dAMD, CSC, DR, GLC, MEM, MYO, RVO, and wAMD—with patient-identity-level splits maintained at approximately 6:2:2 (Wang et al., 2024). Fundus and OCT are explicitly unpaired; the benchmark supports OCT-enhanced disease recognition from fundus images through concept-guided distillation rather than paired multimodal fusion (Wang et al., 2024). This suggests a shift from paired multimodal datasets toward clinically asymmetric training regimes, where expensive structural modalities act as teachers for more ubiquitous imaging.
4. Temporal, infrared, and ultrasound resources
Retinal disease datasets are no longer limited to static optical images. RVD is described as the first publicly available fundus video dataset for vessel segmentation using handheld, smartphone-based acquisition, comprising 635 RGB fundus videos from 415 participants aged 50–75 years across four clinics (Khan et al., 2023). It provides three levels of spatial annotations—binary vessel masks, general artery–vein masks, and fine-grained artery–vein masks—and temporal annotations for spontaneous retinal venous pulsation, including presence, duration, and peak/trough frames (Khan et al., 2023). Because handheld capture introduces eye motion, video jitter, motion blur, nonuniform illumination, and operator variability, the dataset exposes a domain shift relative to bench-top image datasets (Khan et al., 2023).
RAVIR extends retinal vascular analysis into infrared reflectance imaging. It contains 46 Heidelberg Spectralis IR images at 768×768 pixels, 12.5 microns per pixel, with pixel-wise semantic masks for background, vein, and artery classes (Hatamizadeh et al., 2022). The pathology spectrum includes retinal vein occlusion, hypertensive retinopathy, peripapillary atrophy, diabetic retinopathy, isolated vessel tortuosity, high myopia, and media opacities (Hatamizadeh et al., 2022). The dataset was created to support artery–vein segmentation and vessel width measurement, and the SegRAVIR pipeline reports artery Dice 0.8287 and vein Dice 0.8301, alongside artery and vein width estimates of µm and µm respectively (Hatamizadeh et al., 2022).
ERDES moves outside fundus/OCT imaging entirely and addresses retinal detachment in ocular ultrasound. It contains 5,381 deidentified ocular ultrasound video clips from unique participants, collected in the Emergency Department at the University of Arizona, and labeled for Non-RD, PVD, and RD; RD clips are further categorized as Macula_Intact or Macula_Detached, with anatomical subclasses such as TD, ND, and Bilateral (Navard et al., 5 Aug 2025). For the benchmark tasks, PVD is excluded from binary Non-RD vs RD detection. On the test set, 3D U-Net achieved sensitivity 0.950 and accuracy 0.991 for Non-RD vs RD, and sensitivity 0.899 and accuracy 0.882 for Macula-Intact vs Macula-Detached (Navard et al., 5 Aug 2025). A plausible implication is that retinal disease datasets now encode not only pathology labels but also time-critical triage variables.
5. Annotation practice, split strategy, and benchmarking methodology
Annotation protocols vary from image-level diagnosis to pixel-level masks and demographic metadata. OCTDL uses a multi-stage consensus process involving 7 trained medical students, two experienced clinical specialists, and final confirmation by the head of clinic expert (Kulyabin et al., 2023). PALM uses seven ophthalmologists plus one senior ophthalmologist for optic disc masks, fovea points, and lesion masks, with majority voting and senior quality control (Fang et al., 2023). The OCTA DR dataset was graded by two experts—a board-certified ophthalmologist and a retina surgeon—while ERDES used three ocular ultrasound clinical experts plus a fourth expert for quality control and verification (Bidwai et al., 2024, Navard et al., 5 Aug 2025). Several datasets, however, do not report inter-rater agreement statistics.
Split strategy is a recurrent methodological fault line. Patient-level partitioning is explicit in OCTDL, PALM, MultiEYE, and Harvard-GF, and is recommended to avoid leakage across repeated images, visits, or eyes (Kulyabin et al., 2023, Fang et al., 2023, Wang et al., 2024, Luo et al., 2023). By contrast, EyeNet reports a random 70/10/20 split and MuReD reports image-level train/validation organization without patient-level split control (Yang et al., 2018, Rodriguez et al., 2022). HYAMD is longitudinal and includes repeated visits and both eyes, so cross-patient and eye-aware splitting is recommended in its usage guidance (Meisel et al., 7 May 2025). This suggests that reported accuracy is inseparable from split design, especially when datasets include repeated measures or bilateral images.
Benchmark methodology also reflects dataset structure. SynFundus-1M uses 5-fold multi-label stratified cross-validation and constructs a 66-dimensional meta-dataset from out-of-fold predictions of six base models across 11 diseases for XGBoost stacking (Cao-Xue et al., 21 Aug 2025). RFMiD, in a few-shot setting, is used with balanced 5-way, 5-shot episodes and 2 query samples per class, trained on 100 episodes and evaluated on 1,000 test episodes (Khale et al., 4 Dec 2025). SegImgNet applies five-fold cross-validation to AIROGS and e-ROP after resizing images to 256 × 256 pixels, using ROSE oversampling and Weighted Cross-Entropy to handle the severe imbalance of e-ROP (Luo et al., 1 Mar 2025). RetinaVision applies CutMix and MixUp at 224×224, with Adam, categorical cross-entropy, batch size 32, learning rate 0.0001, 50 epochs, and early stopping (Noor et al., 22 Feb 2026).
Evaluation metrics track task type. Classification studies report Accuracy, Precision, Recall, F1, AUC, macro-AUC, and Kappa; segmentation datasets report Dice, IoU, mIoU, mAcc, and mFscore; few-shot and multi-label studies emphasize macro-averaged metrics under imbalance; and fairness benchmarks add DPD, DEO, DEOdds, ES-Acc, and ES-AUC (Cao-Xue et al., 21 Aug 2025, Kulyabin et al., 2023, Luo et al., 2023). Explainability has become part of benchmark design in some OCT and multimodal studies: RetinaVision uses Grad-CAM and LIME, while MultiEYE uses concept bottlenecks and image–concept similarity for OCT-assisted conceptual distillation (Noor et al., 22 Feb 2026, Wang et al., 2024).
6. Generalization, fairness, and recurrent limitations
Several recurring limitations define the current retinal dataset literature. Class imbalance is explicit in RFMiD, OCTDL, e-ROP, ERDES, and C8; MuReD attempts to mitigate this by removing labels with fewer than 30 samples and reassigning them to OTHER, while few-shot work on RFMiD uses balanced episodic sampling and minority-focused augmentation (Khale et al., 4 Dec 2025, Rodriguez et al., 2022, Luo et al., 1 Mar 2025). Exact per-class counts and split ratios are missing in some otherwise widely used resources, including C8 and parts of HYAMD (Noor et al., 22 Feb 2026, Meisel et al., 7 May 2025).
Domain shift is equally pervasive. RetinaVision notes noise and variability from different acquisition systems as an ongoing challenge (Noor et al., 22 Feb 2026). OCTDL emphasizes device specificity, since its scans come from Optovue Avanti whereas many public datasets use Spectralis or Cirrus (Kulyabin et al., 2023). RVD reports substantial gaps between handheld video data and bench-top vessel datasets (Khan et al., 2023). SynFundus-1M demonstrates that models trained exclusively on synthetic data can generalize to real datasets, but with lower external AUC on DR than on AIROGS glaucoma or RFMiD, and explicitly recommends fine-tuning on smaller, expert-labeled real datasets for deployment (Cao-Xue et al., 21 Aug 2025). This suggests that scale alone does not remove modality, device, or site dependence.
Fairness and label semantics are now treated as dataset-level design problems rather than post hoc concerns. Harvard-GF was built as a dedicated fairness dataset with 3,300 subjects, perfectly balanced racial groups, dual-modality imaging, and demographic metadata including race, gender, ethnicity, language, and marital status (Luo et al., 2023). IRFundusSet addresses a different bias source: inconsistency in what public datasets call “healthy,” and therefore introduces a harmonized is_normal label after physical review across ten public cohorts (Githinji et al., 2024). Together, these resources show that retinal disease datasets are increasingly evaluated ոչ only by size or accuracy, but also by whether their labels, splits, and metadata support equitable and clinically interpretable learning.
A final misconception concerns what counts as a “retinal disease dataset.” The literature now includes disease-specific collections, integrated catalogs, synthetic generators, patient-level image–report corpora, cross-modal teacher–student benchmarks, vessel-segmentation resources, and time-resolved video datasets (Du et al., 2024, Githinji et al., 2024, Wang et al., 2024). A plausible implication is that the field has moved from dataset release as isolated curation toward dataset design as experimental infrastructure: label ontology, modality asymmetry, temporal signal, demographic balance, and leakage control are increasingly first-class variables in retinal AI research.