Papers
Topics
Authors
Recent
Search
2000 character limit reached

Evo-RAD: Adaptive Retrieval for Retinal Diagnosis

Updated 6 July 2026
  • Evo-RAD is a retrieval-time adaptation framework that iteratively refines support sets via DELETE, INSERT, and TERMINATE actions for rare retinal disease diagnosis.
  • It employs a graph-based policy to evaluate evidence quality, using both local and global scoring to address the hubness problem.
  • The method optimizes homogeneity and diagnostic performance through GRPO with tailored rewards, achieving significant sensitivity gains over static retrieval methods.

Searching arXiv for Evo-RAD and related retinal diagnosis context. Evo-RAD is a retrieval-time adaptation framework for diagnosing rare retinal diseases from fundus images using a frozen retinal foundation model. It addresses a specific failure mode of conventional retrieval-augmented diagnosis in long-tailed clinical settings: the initial nearest neighbors are often visually similar yet semantically incorrect common diseases, a phenomenon the paper attributes to the hubness problem. Evo-RAD replaces one-shot Top-KK retrieval with a self-evolving agentic procedure in which a graph-based policy iteratively edits the support set through DELETE, INSERT, and TERMINATE actions, and then predicts by majority vote over the evolved references (Xia et al., 22 Jun 2026).

1. Clinical setting and problem formulation

Evo-RAD is designed for rare retinal disease diagnosis under long-tail clinical distributions, where a small number of common diseases dominate the data and many rare diseases have very few examples. In this setting, the paper argues that foundation models can perform well on general screening benchmarks yet still have poor sensitivity for rare conditions, because rare pathologies are underrepresented and can be embedded near common diseases that share superficial visual patterns (Xia et al., 22 Jun 2026).

The method is motivated by limitations in two standard adaptation strategies. First, parameter-efficient fine-tuning can overfit when only a few rare cases are available, or remain dominated by head classes when trained on the full imbalanced dataset. Second, conventional retrieval-augmented diagnosis typically performs a static Top-KK search under a fixed similarity metric. In high-dimensional embedding spaces, some samples become frequent nearest neighbors of many queries; in this application, those hubs are often common-disease cases that repeatedly appear near rare queries. The consequence is a support set polluted by visually similar but clinically incorrect “hard negatives.”

Evo-RAD reframes this as an evidence-selection problem rather than a feature-extraction problem. The central premise is that diagnostic performance on rare diseases can improve if the retrieved evidence is iteratively purified into a more homogeneous and pathologically consistent support set. This suggests that the main target of adaptation is not the frozen backbone itself, but the structure of the reference set presented to the final decision rule.

2. Retrieval as a self-evolving decision process

The paper formulates retrieval as a Markov Decision Process,

M=⟨S,A,T,R⟩,\mathcal{M}=\langle \mathcal{S},\mathcal{A},\mathcal{T},\mathcal{R}\rangle,

with an evolving reference set Xt⊂D\mathcal{X}_t \subset \mathcal{D} that is initialized from standard Top-KK retrieval and then edited sequentially (Xia et al., 22 Jun 2026).

Given a query image qq, a frozen retinal foundation model extracts a visual embedding vqv_q, and the system retrieves the initial reference set

X0⊂D.\mathcal{X}_0 \subset \mathcal{D}.

At each time step, the agent observes a state sts_t summarizing the current support set and chooses from three action types. DELETE removes one current reference case, typically to purge a hub-induced distractor or semantically discordant neighbor. INSERT adds one case from a candidate buffer B⊂D\mathcal{B}\subset\mathcal{D}. The policy does not choose the inserted item directly; instead, the environment inserts

KK0

This means the policy decides when expansion is useful, while the environment chooses the most query-similar unseen candidate. TERMINATE ends refinement when the current set appears sufficiently coherent.

The transition rule is: KK1 Episodes stop when the agent selects TERMINATE or when the step budget is exhausted. In the reported experiments, the policy is allowed at most 10 actions, with a minimum support set size of 2.

This design distinguishes Evo-RAD from static retrieval systems. It also distinguishes it from classifier-centric adaptation: the final diagnosis is not produced by a learned classifier head over the query embedding, but by majority voting over the evolved support set,

KK2

A common misconception is therefore that Evo-RAD is primarily a fine-tuning method; in fact, its defining mechanism is retrieval-time editing of evidence.

3. Graph state representation and policy architecture

Evo-RAD represents the current support set as a graph state

KK3

so that evidence can be evaluated jointly rather than item by item (Xia et al., 22 Jun 2026). This is important because a reference case may appear plausible relative to the query alone while remaining inconsistent with the support set as a whole.

The graph topology is based on semantic consistency among disease labels. Because raw labels may be too coarse, each label KK4 is expanded into fine-grained clinical tags using GPT-5.1,

KK5

and the tag text is encoded with the frozen text encoder of the vision-LLM to obtain a semantic embedding KK6. The adjacency matrix is then defined by cosine similarity,

KK7

Node features are stored in

KK8

with each node represented as

KK9

The statistical descriptor M=⟨S,A,T,R⟩,\mathcal{M}=\langle \mathcal{S},\mathcal{A},\mathcal{T},\mathcal{R}\rangle,0 contains four base metrics and their mean deviations, encoding similarity to the query in visual space, the initial retrieval rank, and agreement with the current set in terms of clinical text and disease categories. The inclusion of mean deviations means the policy observes not only absolute relevance but also deviation from the group consensus.

Graph propagation is performed with a two-layer GCN. With self-loops,

M=⟨S,A,T,R⟩,\mathcal{M}=\langle \mathcal{S},\mathcal{A},\mathcal{T},\mathcal{R}\rangle,1

the update takes the standard normalized form

M=⟨S,A,T,R⟩,\mathcal{M}=\langle \mathcal{S},\mathcal{A},\mathcal{T},\mathcal{R}\rangle,2

After propagation, action scoring combines node-level local scores for deletion candidates with graph-level global scores for insertion and termination: M=⟨S,A,T,R⟩,\mathcal{M}=\langle \mathcal{S},\mathcal{A},\mathcal{T},\mathcal{R}\rangle,3 Pooling is by mean aggregation over nodes.

The architecture therefore encodes a specific inductive bias: evidence quality is a set-level property. DELETE is scored locally because it targets a particular reference, while INSERT and TERMINATE depend on the global coherence of the current set.

4. Learning objective, GRPO optimization, and homogeneity-aware reward

The policy is optimized with Group Relative Policy Optimization (GRPO) rather than an actor-critic scheme (Xia et al., 22 Jun 2026). The paper motivates this choice by noting that rewards are sparse and trajectory quality varies strongly across queries, making critic learning unstable. For each query M=⟨S,A,T,R⟩,\mathcal{M}=\langle \mathcal{S},\mathcal{A},\mathcal{T},\mathcal{R}\rangle,4, a group of trajectories

M=⟨S,A,T,R⟩,\mathcal{M}=\langle \mathcal{S},\mathcal{A},\mathcal{T},\mathcal{R}\rangle,5

is sampled from the old policy, and returns are normalized into relative advantages,

M=⟨S,A,T,R⟩,\mathcal{M}=\langle \mathcal{S},\mathcal{A},\mathcal{T},\mathcal{R}\rangle,6

The policy objective is a GRPO-style relative policy loss with KL regularization: M=⟨S,A,T,R⟩,\mathcal{M}=\langle \mathcal{S},\mathcal{A},\mathcal{T},\mathcal{R}\rangle,7

The reward is explicitly homogeneity-aware. For a trajectory M=⟨S,A,T,R⟩,\mathcal{M}=\langle \mathcal{S},\mathcal{A},\mathcal{T},\mathcal{R}\rangle,8, the total return is

M=⟨S,A,T,R⟩,\mathcal{M}=\langle \mathcal{S},\mathcal{A},\mathcal{T},\mathcal{R}\rangle,9

The trajectory-level reward combines diagnostic correctness, support-set purity, and semantic density. Label purity is

Xt⊂D\mathcal{X}_t \subset \mathcal{D}0

and semantic density is

Xt⊂D\mathcal{X}_t \subset \mathcal{D}1

The terminal reward is

Xt⊂D\mathcal{X}_t \subset \mathcal{D}2

Step-level rewards provide immediate positive feedback for constructive edits: Xt⊂D\mathcal{X}_t \subset \mathcal{D}3 The paper emphasizes a zero-penalty principle: incorrect actions receive reward Xt⊂D\mathcal{X}_t \subset \mathcal{D}4, not a negative penalty. The stated rationale is that negative shaping can hinder early exploration in sparse-reward settings.

A common misunderstanding is that Evo-RAD optimizes retrieval accuracy directly in a static sense. More precisely, it optimizes the homogeneity and diagnostic utility of the final support set under sequential editing. The retrieval target is therefore a coherent evidence cluster rather than merely a similarity-ranked neighbor list.

5. Experimental setup and implementation

Evo-RAD is implemented on top of RetiZero, which remains frozen throughout the method. The trainable component is a small graph policy network with only 116.4K trainable parameters, making the system a lightweight inference-time adapter rather than a full backbone adaptation scheme (Xia et al., 22 Jun 2026).

The paper uses the Retina Image Bank (RIB) rather than heavily used public fundus datasets, on the grounds that strongly exposed benchmarks would obscure genuine generalization. From 30,662 images, a single-label fundus dataset is curated to reduce co-morbidity confounding. Two benchmarks are defined. Rare-20 contains 20 rare diseases, each with Xt⊂D\mathcal{X}_t \subset \mathcal{D}5 images. Retina-31 contains the Rare-20 classes plus 11 common diseases, where common classes satisfy Xt⊂D\mathcal{X}_t \subset \mathcal{D}6 images per class. The train/validation/test split is 70/15/15. For retrieval-based methods, the retrieval corpus is restricted to the training split. The reported split sizes are 524 / 103 / 103 for Rare-20 and 4,737 / 1,001 / 1,023 for Retina-31.

The main baselines fall into three categories. Foundation-model baselines include BiomedCLIP, MedCLIP, EyeCLIP, FLAIR, and RetiZero. PEFT baselines include CoOp, CLIP-Adapter, XCoOp, BiomedCoOp, TDA, Tip-Adapter, and DPC. Retrieval baselines include Static Retrieval and RAC. The core metrics are Accuracy, Macro-F1 score, and Sensitivity, with sensitivity emphasized because the task concerns rare-disease detection under imbalance.

Implementation settings are fixed and compact: candidate buffer size Xt⊂D\mathcal{X}_t \subset \mathcal{D}7, initial retrieval size Xt⊂D\mathcal{X}_t \subset \mathcal{D}8, maximum actions 10, minimum support set size Xt⊂D\mathcal{X}_t \subset \mathcal{D}9, and GRPO group size KK0. Training and evaluation use 3-seed averages on an NVIDIA RTX 4090 (24GB). These choices show that the framework is operationally lightweight relative to full fine-tuning, though inference is necessarily more expensive than static retrieval because the support set is edited over multiple steps.

6. Empirical performance, ablations, and limitations

The paper reports that Evo-RAD improves rare-disease diagnosis substantially, with the abstract stating gains of +21.04% over retinal foundation models and +3.56% over retrieval-based and parameter-efficient fine-tuning methods (Xia et al., 22 Jun 2026). In the Rare-20 benchmark, Evo-RAD reaches 46.28% ACC, 40.99% macro-F1, and 42.43% sensitivity. This exceeds RetiZero zero-shot at 25.24 / 26.37 / 29.35, RetiZero linear probe at 26.86 / 10.15 / 13.05, and RAC at 40.06% sensitivity. On Retina-31, Evo-RAD achieves 65.33% ACC, 51.45% macro-F1, and 49.53% sensitivity, surpassing the reported foundation-model, retrieval, and PEFT baselines.

The empirical case for the method rests especially on the comparison between static retrieval and agentic retrieval. The reported gains indicate that simply retrieving similar examples is not enough; the ability to delete discordant evidence and insert additional support produces a more useful diagnostic set. This is most visible in rare-disease sensitivity, where common-disease hubs are especially damaging.

The ablation studies show that the full design is not reducible to a single component. In the state representation, removing mean deviations reduces sensitivity from 42.43% to 36.21%, while removing base metrics reduces it to 38.41%. This supports the claim that both individual relevance and group-level consensus are needed. For the terminal reward, removing KK1 yields 36.40% sensitivity, removing KK2 yields 33.30%, and removing KK3 yields 36.80%. The largest drop comes from removing purity reward, indicating that explicit support-set purification is central rather than peripheral. Step-level rewards are also important: without KK4, sensitivity falls to 34.85%, and without KK5, to 35.65%. The initial retrieval size is also consequential, with performance peaking at

KK6

which the authors interpret as a balance between insufficient evidence at small KK7 and excessive noise at large KK8.

The method has several important limitations. The paper explicitly states that the current framework is designed for single-label diagnosis and that future work will extend it to multi-label and co-morbidity settings. Additional limitations are implied by the design. Evo-RAD depends on the quality of the initial retrieval pool and candidate buffer, relies on labeled training data for RL reward computation, and constructs its semantic graph from GPT-5.1-generated label expansions. It also does not make the inserted candidate itself a learned action; insertion always follows the fixed visual prior of the most similar unseen candidate in the buffer. A plausible implication is that the framework’s flexibility lies primarily in support-set editing policy rather than in full search over candidate identity.

Its practical significance lies in how it redefines retrieval-augmented diagnosis for rare disease. Evo-RAD does not attempt to solve long-tail bias by changing the frozen representation alone. Instead, it treats evidence acquisition as a sequential decision problem, learns to prefer label-pure and semantically dense support sets, and bases prediction on the resulting evidence cluster. Within retinal diagnosis, that makes it a method for retrieval-time adaptation rather than conventional fine-tuning, and a framework whose central object is the evolving support set rather than the classifier head.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Evo-RAD.