---
title: 'Evo-RAD: Adaptive Retrieval for Retinal Diagnosis'
url: https://www.emergentmind.com/topics/evo-rad
type: topic
---

# Evo-RAD: Adaptive Retrieval for Retinal Diagnosis

Searching arXiv for Evo-RAD and related retinal diagnosis context.
Evo-RAD is a retrieval-time adaptation framework for diagnosing rare retinal diseases from fundus images using a frozen retinal foundation model. It addresses a specific failure mode of conventional retrieval-augmented diagnosis in long-tailed clinical settings: the initial nearest neighbors are often visually similar yet semantically incorrect common diseases, a phenomenon the paper attributes to the hubness problem. Evo-RAD replaces one-shot Top-\(K\) retrieval with a self-evolving agentic procedure in which a graph-based policy iteratively edits the support set through **DELETE**, **INSERT**, and **TERMINATE** actions, and then predicts by majority vote over the evolved references [2606.22955].

## 1. Clinical setting and problem formulation

Evo-RAD is designed for **rare retinal disease diagnosis** under long-tail clinical distributions, where a small number of common diseases dominate the data and many rare diseases have very few examples. In this setting, the paper argues that foundation models can perform well on general screening benchmarks yet still have poor **sensitivity** for rare conditions, because rare pathologies are underrepresented and can be embedded near common diseases that share superficial visual patterns [2606.22955].

The method is motivated by limitations in two standard adaptation strategies. First, parameter-efficient fine-tuning can overfit when only a few rare cases are available, or remain dominated by head classes when trained on the full imbalanced dataset. Second, conventional retrieval-augmented diagnosis typically performs a static Top-\(K\) search under a fixed similarity metric. In high-dimensional embedding spaces, some samples become frequent nearest neighbors of many queries; in this application, those hubs are often common-disease cases that repeatedly appear near rare queries. The consequence is a support set polluted by visually similar but clinically incorrect “hard negatives.”

Evo-RAD reframes this as an evidence-selection problem rather than a feature-extraction problem. The central premise is that diagnostic performance on rare diseases can improve if the retrieved evidence is iteratively purified into a more homogeneous and pathologically consistent support set. This suggests that the main target of adaptation is not the frozen backbone itself, but the structure of the reference set presented to the final decision rule.

## 2. Retrieval as a self-evolving decision process

The paper formulates retrieval as a Markov Decision Process,
\[
\mathcal{M}=\langle \mathcal{S},\mathcal{A},\mathcal{T},\mathcal{R}\rangle,
\]
with an evolving reference set \(\mathcal{X}_t \subset \mathcal{D}\) that is initialized from standard Top-\(K\) retrieval and then edited sequentially [2606.22955].

Given a query image \(q\), a frozen retinal foundation model extracts a visual embedding \(v_q\), and the system retrieves the initial reference set
\[
\mathcal{X}_0 \subset \mathcal{D}.
\]
At each time step, the agent observes a state \(s_t\) summarizing the current support set and chooses from three action types. **DELETE** removes one current reference case, typically to purge a hub-induced distractor or semantically discordant neighbor. **INSERT** adds one case from a candidate buffer \(\mathcal{B}\subset\mathcal{D}\). The policy does not choose the inserted item directly; instead, the environment inserts
\[
x_{\text{ins}}=\arg\max_{x\in \mathcal{B}\setminus \mathcal{X}_t}\mathrm{sim}(q,x).
\]
This means the policy decides when expansion is useful, while the environment chooses the most query-similar unseen candidate. **TERMINATE** ends refinement when the current set appears sufficiently coherent.

The transition rule is:
\[
\mathcal{X}_{t+1}= \begin{cases} 
\mathcal{X}_t\setminus\{x\}, & a_t\in\mathcal{A}_{\text{del}},\\
\mathcal{X}_t\cup\{x_{\text{ins}}\}, & a_t\in\mathcal{A}_{\text{ins}},\\
\mathcal{X}_t, & a_t=a_{\text{term}}.
\end{cases}
\]
Episodes stop when the agent selects **TERMINATE** or when the step budget is exhausted. In the reported experiments, the policy is allowed at most **10 actions**, with a minimum support set size of **2**.

This design distinguishes Evo-RAD from static retrieval systems. It also distinguishes it from classifier-centric adaptation: the final diagnosis is not produced by a learned classifier head over the query embedding, but by majority voting over the evolved support set,
\[
\hat{y} =\argmax_{c\in\mathcal{Y}} \sum_{x\in\mathcal{X}_T}\mathbb{I}(y_x=c).
\]
A common misconception is therefore that Evo-RAD is primarily a fine-tuning method; in fact, its defining mechanism is retrieval-time editing of evidence.

## 3. Graph state representation and policy architecture

Evo-RAD represents the current support set as a graph state
\[
s_t=(\mathcal{G}_t,\mathbf{H}_t),
\]
so that evidence can be evaluated jointly rather than item by item [2606.22955]. This is important because a reference case may appear plausible relative to the query alone while remaining inconsistent with the support set as a whole.

The graph topology is based on semantic consistency among disease labels. Because raw labels may be too coarse, each label \(y\in\mathcal{Y}\) is expanded into fine-grained clinical tags using **GPT-5.1**,
\[
\mathrm{Tags}(y),
\]
and the tag text is encoded with the frozen text encoder of the vision-language model to obtain a semantic embedding \(u_x\). The adjacency matrix is then defined by cosine similarity,
\[
(\mathbf{A}_t)_{ij} := \frac{u_{x_i}^\top u_{x_j}}{\|u_{x_i}\|_2\,\|u_{x_j}\|_2}, \quad x_i,x_j\in\mathcal{X}_t.
\]

Node features are stored in
\[
\mathbf{H}_t \in \mathbb{R}^{|\mathcal{X}_t| \times D},
\]
with each node represented as
\[
h(x) = [\, v_x \| \phi_{\text{stat}}(x, \mathcal{X}_t, q) \,].
\]
The statistical descriptor \(\phi_{\text{stat}}\in\mathbb{R}^8\) contains **four base metrics and their mean deviations**, encoding similarity to the query in visual space, the initial retrieval rank, and agreement with the current set in terms of clinical text and disease categories. The inclusion of mean deviations means the policy observes not only absolute relevance but also deviation from the group consensus.

Graph propagation is performed with a two-layer GCN. With self-loops,
\[
\hat{\mathbf{A}}_t=\mathbf{A}_t+\mathbf{I},
\]
the update takes the standard normalized form
\[
\mathbf{H}_t^{(l+1)}= \sigma\!\left( \hat{\mathbf{D}}_t^{-\frac{1}{2}} \hat{\mathbf{A}}_t \hat{\mathbf{D}}_t^{-\frac{1}{2}} \mathbf{H}_t^{(l)}\mathbf{W}^{(l)} \right),
\quad \mathbf{H}_t^{(0)}=\mathbf{H}_t.
\]
After propagation, action scoring combines node-level local scores for deletion candidates with graph-level global scores for insertion and termination:
\[
\pi_\theta(a_t\mid s_t)=\operatorname{Softmax}\!\Big( \Big[ \{\operatorname{MLP}_{\text{local}}(h_x^{(L)})\}_{x\in\mathcal{X}_t},\; \operatorname{MLP}_{\text{global}}(\operatorname{Pool}(\mathbf{H}_t^{(L)})) \Big] \Big)_{a_t}.
\]
Pooling is by mean aggregation over nodes.

The architecture therefore encodes a specific inductive bias: evidence quality is a set-level property. DELETE is scored locally because it targets a particular reference, while INSERT and TERMINATE depend on the global coherence of the current set.

## 4. Learning objective, GRPO optimization, and homogeneity-aware reward

The policy is optimized with **Group Relative Policy Optimization (GRPO)** rather than an actor-critic scheme [2606.22955]. The paper motivates this choice by noting that rewards are sparse and trajectory quality varies strongly across queries, making critic learning unstable. For each query \(q\), a group of trajectories
\[
\mathcal{T}_G=\{\tau_1,\ldots,\tau_G\}
\]
is sampled from the old policy, and returns are normalized into relative advantages,
\[
\hat{A}_i= \frac{R(\tau_i)-\operatorname{Mean}(\{R(\tau_j)\}_{j=1}^G)}
{\operatorname{Std}(\{R(\tau_j)\}_{j=1}^G)+\epsilon}.
\]
The policy objective is a GRPO-style relative policy loss with KL regularization:
\[
\mathcal{J}(\theta)= \mathbb{E}_{q}\left[ \frac{1}{G}\sum_{i=1}^G \left( \frac{\pi_{\theta}(\tau_i)}{\pi_{\theta_{\text{old}}}(\tau_i)}\hat{A}_i -\beta\,\mathbb{D}_{\text{KL}}\!\left(\pi_{\theta}(\cdot\mid q)\,\|\,\pi_{\text{ref}}(\cdot\mid q)\right) \right) \right].
\]

The reward is explicitly **homogeneity-aware**. For a trajectory \(\tau\), the total return is
\[
R(\tau)= R_{\text{traj}}(\mathcal{X}_T) + \sum_{t=0}^{T-1} r_{\text{step}}(s_t,a_t).
\]
The trajectory-level reward combines diagnostic correctness, support-set purity, and semantic density. Label purity is
\[
\text{Pur}(\mathcal{X}) = \frac{1}{|\mathcal{X}|}\sum_{x\in\mathcal{X}}\mathbb{I}[y_x=y^\star],
\]
and semantic density is
\[
\text{Den}(\mathcal{X}) = \frac{2}{|\mathcal{X}|(|\mathcal{X}|-1)}\sum_{i<j}\cos(u_{x_i},u_{x_j}).
\]
The terminal reward is
\[
R_{\text{traj}}(\mathcal{X}_T) =
\underbrace{\alpha\,\mathbb{I}\!\left[\hat{y}(\mathcal{X}_T)=y^\star\right]}_{R_{\text{acc}}}
+\underbrace{\beta\Big(\text{Pur}(\mathcal{X}_T)-\text{Pur}(\mathcal{X}_0)\Big)}_{R_{\text{purity}}}
+\underbrace{\gamma\,\text{Den}(\mathcal{X}_T)}_{R_{\text{density}}}.
\]

Step-level rewards provide immediate positive feedback for constructive edits:
\[
r_{\text{step}}(s_t,a_t)=
\underbrace{\eta_{\text{ins}}\;\mathbb{I}\!\left[a_t\in\mathcal{A}_{\text{ins}}\right]\mathbb{I}\!\left[y_{x_{\text{ins}}}=y^\star\right]}_{r_{\text{ins}}}
+
\underbrace{\eta_{\text{del}}\;\mathbb{I}\!\left[a_t\in\mathcal{A}_{\text{del}}\right]\mathbb{I}\!\left[y_{x_{\text{del}}}\neq y^\star\right]}_{r_{\text{del}}}.
\]
The paper emphasizes a **zero-penalty principle**: incorrect actions receive reward \(0\), not a negative penalty. The stated rationale is that negative shaping can hinder early exploration in sparse-reward settings.

A common misunderstanding is that Evo-RAD optimizes retrieval accuracy directly in a static sense. More precisely, it optimizes the homogeneity and diagnostic utility of the final support set under sequential editing. The retrieval target is therefore a coherent evidence cluster rather than merely a similarity-ranked neighbor list.

## 5. Experimental setup and implementation

Evo-RAD is implemented on top of **RetiZero**, which remains frozen throughout the method. The trainable component is a small graph policy network with only **116.4K trainable parameters**, making the system a lightweight inference-time adapter rather than a full backbone adaptation scheme [2606.22955].

The paper uses the **Retina Image Bank (RIB)** rather than heavily used public fundus datasets, on the grounds that strongly exposed benchmarks would obscure genuine generalization. From 30,662 images, a single-label fundus dataset is curated to reduce co-morbidity confounding. Two benchmarks are defined. **Rare-20** contains 20 rare diseases, each with \(15 < N \le 75\) images. **Retina-31** contains the Rare-20 classes plus 11 common diseases, where common classes satisfy \(N > 200\) images per class. The train/validation/test split is **70/15/15**. For retrieval-based methods, the retrieval corpus is restricted to the training split. The reported split sizes are **524 / 103 / 103** for Rare-20 and **4,737 / 1,001 / 1,023** for Retina-31.

The main baselines fall into three categories. Foundation-model baselines include **BiomedCLIP**, **MedCLIP**, **EyeCLIP**, **FLAIR**, and **RetiZero**. PEFT baselines include **CoOp**, **CLIP-Adapter**, **XCoOp**, **BiomedCoOp**, **TDA**, **Tip-Adapter**, and **DPC**. Retrieval baselines include **Static Retrieval** and **RAC**. The core metrics are **Accuracy**, **Macro-F1 score**, and **Sensitivity**, with sensitivity emphasized because the task concerns rare-disease detection under imbalance.

Implementation settings are fixed and compact: candidate buffer size \(|\mathcal{B}|=100\), initial retrieval size \(K=8\), maximum actions 10, minimum support set size \(min\_size=2\), and GRPO group size \(G=8\). Training and evaluation use **3-seed averages** on an **NVIDIA RTX 4090 (24GB)**. These choices show that the framework is operationally lightweight relative to full fine-tuning, though inference is necessarily more expensive than static retrieval because the support set is edited over multiple steps.

## 6. Empirical performance, ablations, and limitations

The paper reports that Evo-RAD improves rare-disease diagnosis substantially, with the abstract stating gains of **+21.04%** over retinal foundation models and **+3.56%** over retrieval-based and parameter-efficient fine-tuning methods [2606.22955]. In the Rare-20 benchmark, Evo-RAD reaches **46.28% ACC**, **40.99% macro-F1**, and **42.43% sensitivity**. This exceeds **RetiZero zero-shot** at **25.24 / 26.37 / 29.35**, **RetiZero linear probe** at **26.86 / 10.15 / 13.05**, and **RAC** at **40.06% sensitivity**. On Retina-31, Evo-RAD achieves **65.33% ACC**, **51.45% macro-F1**, and **49.53% sensitivity**, surpassing the reported foundation-model, retrieval, and PEFT baselines.

The empirical case for the method rests especially on the comparison between static retrieval and agentic retrieval. The reported gains indicate that simply retrieving similar examples is not enough; the ability to delete discordant evidence and insert additional support produces a more useful diagnostic set. This is most visible in rare-disease sensitivity, where common-disease hubs are especially damaging.

The ablation studies show that the full design is not reducible to a single component. In the state representation, removing mean deviations reduces sensitivity from **42.43%** to **36.21%**, while removing base metrics reduces it to **38.41%**. This supports the claim that both individual relevance and group-level consensus are needed. For the terminal reward, removing \(R_{\text{acc}}\) yields **36.40%** sensitivity, removing \(R_{\text{purity}}\) yields **33.30%**, and removing \(R_{\text{density}}\) yields **36.80%**. The largest drop comes from removing purity reward, indicating that explicit support-set purification is central rather than peripheral. Step-level rewards are also important: without \(r_{\text{ins}}\), sensitivity falls to **34.85%**, and without \(r_{\text{del}}\), to **35.65%**. The initial retrieval size is also consequential, with performance peaking at
\[
K=8,
\]
which the authors interpret as a balance between insufficient evidence at small \(K\) and excessive noise at large \(K\).

The method has several important limitations. The paper explicitly states that the current framework is designed for **single-label diagnosis** and that future work will extend it to **multi-label and co-morbidity settings**. Additional limitations are implied by the design. Evo-RAD depends on the quality of the initial retrieval pool and candidate buffer, relies on labeled training data for RL reward computation, and constructs its semantic graph from **GPT-5.1**-generated label expansions. It also does not make the inserted candidate itself a learned action; insertion always follows the fixed visual prior of the most similar unseen candidate in the buffer. A plausible implication is that the framework’s flexibility lies primarily in support-set editing policy rather than in full search over candidate identity.

Its practical significance lies in how it redefines retrieval-augmented diagnosis for rare disease. Evo-RAD does not attempt to solve long-tail bias by changing the frozen representation alone. Instead, it treats evidence acquisition as a sequential decision problem, learns to prefer label-pure and semantically dense support sets, and bases prediction on the resulting evidence cluster. Within retinal diagnosis, that makes it a method for retrieval-time adaptation rather than conventional fine-tuning, and a framework whose central object is the evolving support set rather than the classifier head.

Source: https://www.emergentmind.com/topics/evo-rad