---
title: 'EYE4ALL: Accessibility & OCT Vision Dataset'
url: https://www.emergentmind.com/topics/eye4all
type: topic
---

# EYE4ALL: Accessibility & OCT Vision Dataset

Searching arXiv for the named EYE4ALL papers and closely related supporting work.
EYE4ALL is a name used for multiple research artifacts spanning accessibility-oriented vision-language evaluation, ophthalmic imaging, and, by extension in related literature, enabling technologies for assistive vision, eye tracking, and compact near-eye optics. In the most explicit recent usage, EYE4ALL denotes a blind/low-vision-centered text-image-to-text benchmark for evaluating large vision-language model responses in navigation assistance [2510.00766]. In another usage, it denotes an open-source whole-eye optical coherence tomography dataset built from a spectrally-multiplexed whole-eye OCT system for anterior segment segmentation, ocular biometry, and personalized eye modeling [2605.19191]. Related work situates these EYE4ALL usages within a broader technical ecosystem that includes wearable assistive perception [2303.13863], large-scale eye-image annotation for gaze and eye movement analysis [2102.02115], robust multi-view eye tracking [1711.05444], holographic cataract and intraocular lens simulation [2304.00548], and wide-field meta-optic eyepieces for near-eye systems [2406.14725].

## 1. EYE4ALL as an accessibility benchmark for text-image-to-text evaluation

In the vision-language setting, EYE4ALL is a benchmark explicitly grounded in accessibility and blind/low-vision needs. It is built around text-image-to-text responses: an image plus a scene-relevant request are given to a large vision-language model, and the resulting response is judged for how well it supports a blind or low-vision user navigating a real scene [2510.00766]. The benchmark is intended to capture not merely image-text consistency, but whether a response is useful, safe, sufficiently detailed, concise, and free of hallucinations for blind/low-vision-oriented assistance.

The dataset is divided into two components. **EYE4ALLPref** is a pairwise preference dataset with chosen/rejected comparisons. **EYE4ALLMulti** is a pointwise multi-objective dataset with fine-grained human scores across seven dimensions: **Direction Accuracy**, **Depth Accuracy**, **Safety**, **Sufficiency**, **Conciseness**, **Hallucination**, and **Overall Quality** [2510.00766]. The benchmark is constructed from pedestrian and navigation scenes drawn from the **Sideguide** and **Sidewalk** corpora, paired with requests described as plausible for blind/low-vision navigation contexts.

The generation pipeline begins with scene images, pairs them with a request, and then produces responses from **Qwen2-VL**, **LLaVA-1.6**, and **InternLM-XComposer2-VL (InternLM-X2-VL)**. The models are prompted to answer from the perspective of blind or low-vision users. For each instance, one of the three model responses is selected randomly and then refined by **GPT-4o mini** using a blind/low-vision-oriented prompt [2510.00766]. This makes EYE4ALL a curated, model-generated-and-refined benchmark rather than a corpus of purely human-authored responses.

The human annotation protocol is central to the benchmark’s design. The judgments were collected from **25 sighted human annotators**, approved by the **IRB**. Each annotator completed **100 items** online, taking about **2 to 3 hours**. The guideline emphasized scene description, distance and direction to the target, obstacles to watch for, and step-by-step directions. The paper reports an **average annotator agreement of 33.21** with a standard deviation of **17.70** [2510.00766]. Non-hallucination dimensions are averaged from **2–3 human judgment scores on a 1–5 Likert scale**, while **Hallucination** is binary, with **0 = presence** and **1 = absence**.

A distinctive feature of EYE4ALL is the separation of spatial guidance into direction and depth. The paper notes that **Direction** and **Depth** are separated because precise spatial guidance is especially important for blind or low-vision users [2510.00766]. This reflects a benchmark design oriented toward embodied navigation rather than generic captioning quality.

## 2. EYE4ALL within multi-objective alignment modeling

EYE4ALL serves as the main accessibility-oriented evaluation target for **MULTI-TAP (Multi-Objective Task-Aware Predictor)**, a plug-and-play architecture for single-objective and multi-objective image-text alignment prediction [2510.00766]. In this context, EYE4ALL is not only a benchmark but also a training and evaluation substrate for human-aligned reward modeling in assistive scenarios.

MULTI-TAP is trained in two stages. In the first stage, a reward head predicts a single scalar alignment score from frozen or fine-tuned large vision-language model hidden states using mean squared error:
$$
\min_\theta \sum^N_{i} (r_i - h_i)^2
$$
where \(r_i\) is the predicted scalar score, \(h_i\) is the human judgment score, and \(\theta\) denotes the parameters of the large vision-language model and reward head [2510.00766]. In the second stage, a frozen large vision-language model embedding \(z_i\) is input to a ridge regression model to predict a vector of human-interpretable dimension scores:
$$
\hat y_i = W z_i + b
$$
with training objective
$$
\min_{W,b}\ \sum_{i=1}^{N}\bigl\|y_i - W z_i - b\bigr\|_2^2 + \alpha\|W\|_F^2 .
$$

The paper explicitly states that MULTI-TAP does **not** aggregate multi-objective outputs into a single score. Instead, the **overall score** comes from the first-stage reward head, while the **dimension-specific scores** come from the second-stage ridge regression head [2510.00766]. This design is contrasted with models such as **VisionREW-S**, which aggregate multi-objective outputs with learned weights.

On **EYE4ALLMulti**, the strongest version of MULTI-TAP reaches about **52.08**, whereas **VisionREW-S** achieves about **47.63** [2510.00766]. The paper also emphasizes computational efficiency: **VisionREW-S** is reported to require about **51 days** on a single RTX A6000 for the VisionREW setting, while MULTI-TAP completes training and inference in about **4 hours** for a Qwen2-based variant and **11 hours** for a LLaMA-3.2-based variant. A plausible implication is that EYE4ALL was designed not only as an evaluation instrument but also as a stress test for practical, interpretable, and efficient assistive-model assessment.

## 3. EYE4ALL as an open-source whole-eye OCT dataset

A second, unrelated but technically significant usage of the name EYE4ALL appears in ophthalmic imaging. Here, EYE4ALL denotes an open-source dataset derived from a **spectrally-multiplexed whole-eye OCT (WEOCT)** system, intended to fill a gap in public data for **anterior segment segmentation**, **ocular biometry**, and **whole-eye reconstruction** [2605.19191]. The release contains **6,621 processed volumes** from **276 unique participants**, together with segmentation and calibrated 3D anterior point clouds.

The imaging system employs two synchronized swept sources: a **1310 nm** anterior-segment channel with **60 nm** bandwidth and **200 kHz** sweep rate, and a **1060/1064 nm** retinal channel with **102 nm** bandwidth and **200 kHz** sweep rate [2605.19191]. The acquisition protocol uses
$$
500 \text{ A-scans/B-scan}, \quad 250 \text{ B-scans/volume}
$$
at a volumetric rate of about **1.6 Hz**. The anterior channel covers roughly
$$
35 \text{ mm} \times 17.5 \text{ mm},
$$
while the retinal channel covers about
$$
36^\circ \times 18^\circ.
$$

The dataset includes three categories of released material: raw volumes, processed data, and manual annotations. Raw volumes comprise spectral OCT B-scan stacks and scan waveforms. Processed data include B-scans, segmentation masks, boundary pixels, full calibrated point clouds, and segmentation labels attached to the point clouds. Manual annotations are provided for a subset from **34 participants**, with annotations available for each B-scan [2605.19191].

The processing pipeline has four stages: **deep learning-based surface segmentation**, **3D distortion correction**, **surface fitting**, and **ray-tracing refraction correction**. The segmentation network is a **DenseUNet** trained on **51,398 annotated B-scans** from **552 volumes** from **30 subjects**, with a validation set of **7,951 B-scans** from **79 volumes** from **8 subjects** [2605.19191]. The annotations label **cornea/sclera** and **iris**, with cornea and sclera grouped during annotation because their boundary is ambiguous in OCT.

Distortion calibration uses a custom **dot-pattern target**, a **hexapod robot**, multiple known **Z positions**, and per-pixel ray recovery, followed by **3rd-order 2D polynomial fits** that map galvanometer voltages to physical 3D rays [2605.19191]. After surface fitting, ray-tracing refraction correction accounts for the fact that OCT measures optical path length rather than true geometric depth. The pipeline uses fixed refractive indices
$$
n_{\text{cornea}} = 1.376, \qquad n_{\text{aqueous}} = 1.336
$$
and applies Snell’s law,
$$
n_1 \sin \theta_1 = n_2 \sin \theta_2,
$$
at the relevant interfaces.

Validation is reported through both phantom and human studies. On a **12.5 mm radius ceramic sphere**, the paper reports a **median radius error of 14 μm**, **p95 radius error of 19 μm**, **median 3D center localization error of 30 μm**, and **p95 of 33 μm** [2605.19191]. In the phantom pupil-center experiment, the **median error** is **22 μm** and **p95** is **65 μm**. The human study framework includes **300+ participants** in the abstract framing, with **276 unique participants** in the released QA-passed dataset.

## 4. Annotation, geometry, and eye-modeling context

The biomedical EYE4ALL release belongs to a broader trend toward geometry-rich eye datasets. A key comparator is **TEyeD**, which provides **over 20 million** real-world eye images collected with **seven head-mounted eye trackers**, including devices integrated into **VR** and **AR** systems [2102.02115]. TEyeD includes **2D and 3D landmarks**, **semantic segmentations**, **3D eyeball annotation**, **gaze vectors**, and **eye movement types**.

TEyeD’s annotations span **pupil**, **iris**, and **eyelids** in both 2D and 3D, plus gaze and movement labels such as fixations, saccades, smooth pursuits, blinks, and invalid/error cases [2102.02115]. The dataset uses a semi-supervised **Multiple Annotation Maturation (MAM)** workflow, beginning with **20,000 manually annotated images**, and iteratively refining labels with a **ResNet-50** landmark regressor and a residual **U-Net** segmentation model. Images are flagged for manual review when the Jaccard agreement between derived segmentations falls below **0.9**, using
$$
\frac{ \text{ResNet50} \cap \text{U-Net} }{ \text{ResNet50} \cup \text{U-Net} }.
$$

A notable point in TEyeD is the distinction between the pupil’s image-plane appearance and its physical 3D position. Because of corneal refraction and steep camera angles, the pupil often appears displaced relative to the iris center; accordingly, the dataset adjusts 3D landmarks and 3D segmentation to the iris center rather than directly using the observed pupil appearance [2102.02115]. This geometry-aware treatment complements the biometric goals of the OCT-based EYE4ALL release. This suggests that the two datasets address adjacent but distinct layers of eye modeling: TEyeD is image-centric and behavior-centric, whereas the OCT EYE4ALL is anatomy-centric and biometry-centric.

The relevance to personalized eye modeling is explicit in the OCT work, which states applications including **visual/optical axis measurement**, **eye tracking model calibration**, **VR/AR/MR eye modeling**, and **personalized gaze estimation** [2605.19191]. In that sense, EYE4ALL as whole-eye OCT infrastructure contributes anatomical priors and calibrated geometry that can underpin future gaze-estimation and XR calibration pipelines.

## 5. Assistive-wearable interpretations and related perceptual systems

Although not itself named EYE4ALL, **MagicEye** provides a closely related assistive-wearable model that helps situate accessibility-oriented interpretations of the term. MagicEye is an **intelligent wearable eyeglass-based system** for visually impaired users that integrates **real-time object detection**, **facial recognition**, **currency recognition**, **GPS-based navigation**, and **proximity-based obstacle alerts** into one device [2303.13863].

Its object detection core is based on **YOLOv5**, with a **CSPNet backbone**, **PANet**, and a YOLO detection head that outputs **20×20**, **40×40**, and **80×80** feature maps [2303.13863]. The detector is trained on **35 classes** selected from **Open Images Dataset V4**, using images resized to **416 × 416**, **batch size of 32**, and **SGD with momentum**. The paper reports **mAP = 68.2%** on **1661** held-out samples, using a training set of **over 13,200 samples**.

The paper also reports distinct subsystem metrics. **Currency identification** reaches **99.75% accuracy**, **99.86% F1-score**, **100% recall**, and **99.72% precision**. **Facial recognition** reaches **94.5% accuracy**, **94.21% F1-score**, **89.60% recall**, and **99.34% precision** [2303.13863]. Face recognition uses an ensemble of **FaceNet** and **VGG-Face**, with **MTCNN** for face detection and cosine similarity
$$
\text{Cosine Similarity} = \frac{A \cdot B}{\|A\|\|B\|}
$$
for matching.

MagicEye is architecturally relevant because it frames assistive vision as a multi-modal always-available wearable. Its pipeline begins when the user activates the device or the proximity sensor triggers it, followed by image capture, AI processing, object detection, audio feedback, optional face and currency identification, GPS-based route guidance, and immediate obstacle warnings [2303.13863]. This suggests one plausible practical interpretation of EYE4ALL as a broader research direction: a convergence of multimodal sensing, model-based perception, and accessibility-focused output, even when the name itself is used more narrowly in individual papers.

## 6. Eye tracking, near-eye optics, and perceptual display infrastructure

The EYE4ALL-related literature also includes enabling hardware and algorithms for robust eye sensing and near-eye image delivery. In gaze tracking, **Robust Real-Time Multi-View Eye Tracking** introduces a multi-camera framework that simultaneously acquires multiple eye appearances and fuses the resulting gaze outputs through adaptive weighting [1711.05444]. The prototype uses **three synchronized PointGrey Flea3 monochrome cameras**, each with **1280×1024** resolution, **8 mm** lenses, and **850 nm** near-infrared LEDs, and runs at **30 fps**.

The system computes up to \(2C\) gaze outputs per frame for \(C\) cameras and fuses them as
$$
{\bf z^{*} = \sum_c \sum_e {\bf z_c^e}w_c^e, \ \sum_c \sum_e w_c^e =1, \hspace{2mm}  e \in\{L,R\},~\hspace{2mm} c \in\{1,2,..,C\}, 
$$
with head-pose-based or gazing-behavior-based weights [1711.05444]. In the reported experiments, the multi-view system achieves about **1 degree accuracy** under challenging scenarios and nearly **100% availability**, including improvements for users with glasses.

At a lower-level ocular-feature scale, **Custom Video-Oculography Device and Its Application to Fourth Purkinje Image Detection during Saccades** presents a full-resolution **MJPEG** video-oculography platform supporting offline reanalysis of **pupil**, **First Purkinje image (P1)**, and **Fourth Purkinje image (P4)** detection [1904.07361]. The system saves every frame at full resolution, supports recordings up to **500 fps**, and uses a staged blob-selection procedure for P4 detection inside a pupil-centered area of interest. Candidate P4 blobs must have area between **5 and 30 pixels**, maximum brightness above an adaptive threshold, and are selected by proximity to the pupil center. The paper shows that a **P1 − P4** signal qualitatively resembles historical dual-Purkinje-image tracking during saccades [1904.07361].

On the display side, **Wide Field of View Large Aperture Meta-Doublet Eyepiece** demonstrates a wide-field meta-optic eyepiece relevant to compact near-eye systems [2406.14725]. The paper reports a meta-doublet eyepiece with **greater than 60°** field of view and an entrance aperture of **2.1 cm**, designed at **633 nm**. For the **2 cm meta-doublet**, reported specifications include **21.0 mm** entrance aperture, **60° full FoV**, **5.4 mm** pupil diameter, **15.0 mm** eye relief, **15.17 mm** effective focal length, **0.18** numerical aperture, and **35.7 mm** total track length [2406.14725]. A comparable refractive triplet is listed at **43.0 mm** total track length and **45°** apparent FoV.

The meta-doublet consists of periodic arrays of **square silicon nitride pillars** on quartz, with **750 nm** tall pillars, widths from **80 to 270 nm**, **350 nm** pitch, **SiN \(n = 2.04\)**, and quartz substrate **\(n = 1.46\)** [2406.14725]. Phase delays are computed with **rigorous coupled-wave analysis (RCWA)** and optimized in **Zemax OpticStudio**. The first metasurface acts as an entrance aperture and corrective plate, while the second provides most of the focusing power. For near-eye accessibility devices, a plausible implication is that such optics could eventually support thinner and lighter monocular or binocular perceptual-assistance displays.

A related perceptual-simulation direction appears in **Artificial Eye Model and Holographic Display Based IOL Simulator**, which combines an artificial pseudophakic eye with a holographic simulator for preoperative IOL counseling [2304.00548]. The benchtop system uses a **phase-only spatial light modulator (LCoS SLM)** with **1920 × 1080** resolution, **8 µm** pixel pitch, and **60 fps** phase modulation, together with **He-Ne laser, wavelength 632.9 nm**, and **LED, wavelength 530 nm** illumination [2304.00548]. It evaluates a **monofocal IOL: 14.5 D** and a **bifocal IOL: 15.5 + 3.25 D**, showing halos, glare, contrast loss, and coherence-dependent interference fringes under laser illumination.

## 7. Conceptual scope, limitations, and recurring themes

Across the cited literature, EYE4ALL does not denote a single unified research program. Rather, it names at least two specific resources in different subfields: an accessibility-centered text-image-to-text benchmark [2510.00766] and an OCT-based whole-eye imaging dataset [2605.19191]. The surrounding literature shows recurring concerns with robustness, human-centered evaluation, geometry-rich annotation, and deployable sensing or display hardware.

In the accessibility benchmark usage, the central concern is whether large vision-language model outputs are safe, sufficient, spatially accurate, concise, and non-hallucinatory for blind or low-vision navigation assistance [2510.00766]. In the OCT usage, the emphasis is on anatomically accurate 3D reconstruction, segmentation, and biometry, enabled by spectrally multiplexed acquisition and physics-based correction [2605.19191]. These are distinct technical agendas, but both are organized around making eye- and vision-related systems more practically useful.

Several limitations are explicit in the source papers. The benchmark paper states that even strong models remain limited on challenging multi-objective tasks such as EYE4ALL, and that hallucination-free, human-preferred blind/low-vision-oriented generation remains open [2510.00766]. The OCT dataset paper notes substantial scan rejection from motion and poor retinal visibility, with about **45% of collected volumes** acceptable after QA [2605.19191]. MagicEye leaves long-term battery life, full embedded deployment constraints, and failure modes under poor lighting or crowded scenes largely unexamined [2303.13863]. TEyeD includes invalid and no-eye frames and depends partly on geometrically derived annotations [2102.02115]. The meta-optic eyepiece remains monochromatic at **633 nm**, which constrains it mainly to monochrome near-eye display and night-vision contexts [2406.14725].

Taken together, the literature indicates that EYE4ALL functions less as a singular technical object than as a recurring label at the intersection of accessibility, ophthalmic measurement, and eye-centered computational imaging. One branch evaluates whether multimodal AI can assist blind or low-vision users safely and effectively in real scenes [2510.00766]. Another provides calibrated anatomical data for ocular reconstruction and biometry [2605.19191]. Surrounding work on wearable assistance [2303.13863], eye-image annotation [2102.02115], multi-view tracking [1711.05444], Purkinje-based oculography [1904.07361], holographic ocular simulation [2304.00548], and meta-optic eyepieces [2406.14725] defines the broader technical context in which EYE4ALL-related systems may develop.

Source: https://www.emergentmind.com/topics/eye4all