Evo-RAD: Adaptive Retrieval for Retinal Diagnosis
- Evo-RAD is a retrieval-time adaptation framework that iteratively refines support sets via DELETE, INSERT, and TERMINATE actions for rare retinal disease diagnosis.
- It employs a graph-based policy to evaluate evidence quality, using both local and global scoring to address the hubness problem.
- The method optimizes homogeneity and diagnostic performance through GRPO with tailored rewards, achieving significant sensitivity gains over static retrieval methods.
Searching arXiv for Evo-RAD and related retinal diagnosis context. Evo-RAD is a retrieval-time adaptation framework for diagnosing rare retinal diseases from fundus images using a frozen retinal foundation model. It addresses a specific failure mode of conventional retrieval-augmented diagnosis in long-tailed clinical settings: the initial nearest neighbors are often visually similar yet semantically incorrect common diseases, a phenomenon the paper attributes to the hubness problem. Evo-RAD replaces one-shot Top- retrieval with a self-evolving agentic procedure in which a graph-based policy iteratively edits the support set through DELETE, INSERT, and TERMINATE actions, and then predicts by majority vote over the evolved references (Xia et al., 22 Jun 2026).
1. Clinical setting and problem formulation
Evo-RAD is designed for rare retinal disease diagnosis under long-tail clinical distributions, where a small number of common diseases dominate the data and many rare diseases have very few examples. In this setting, the paper argues that foundation models can perform well on general screening benchmarks yet still have poor sensitivity for rare conditions, because rare pathologies are underrepresented and can be embedded near common diseases that share superficial visual patterns (Xia et al., 22 Jun 2026).
The method is motivated by limitations in two standard adaptation strategies. First, parameter-efficient fine-tuning can overfit when only a few rare cases are available, or remain dominated by head classes when trained on the full imbalanced dataset. Second, conventional retrieval-augmented diagnosis typically performs a static Top- search under a fixed similarity metric. In high-dimensional embedding spaces, some samples become frequent nearest neighbors of many queries; in this application, those hubs are often common-disease cases that repeatedly appear near rare queries. The consequence is a support set polluted by visually similar but clinically incorrect “hard negatives.”
Evo-RAD reframes this as an evidence-selection problem rather than a feature-extraction problem. The central premise is that diagnostic performance on rare diseases can improve if the retrieved evidence is iteratively purified into a more homogeneous and pathologically consistent support set. This suggests that the main target of adaptation is not the frozen backbone itself, but the structure of the reference set presented to the final decision rule.
2. Retrieval as a self-evolving decision process
The paper formulates retrieval as a Markov Decision Process,
with an evolving reference set that is initialized from standard Top- retrieval and then edited sequentially (Xia et al., 22 Jun 2026).
Given a query image , a frozen retinal foundation model extracts a visual embedding , and the system retrieves the initial reference set
At each time step, the agent observes a state summarizing the current support set and chooses from three action types. DELETE removes one current reference case, typically to purge a hub-induced distractor or semantically discordant neighbor. INSERT adds one case from a candidate buffer . The policy does not choose the inserted item directly; instead, the environment inserts
0
This means the policy decides when expansion is useful, while the environment chooses the most query-similar unseen candidate. TERMINATE ends refinement when the current set appears sufficiently coherent.
The transition rule is: 1 Episodes stop when the agent selects TERMINATE or when the step budget is exhausted. In the reported experiments, the policy is allowed at most 10 actions, with a minimum support set size of 2.
This design distinguishes Evo-RAD from static retrieval systems. It also distinguishes it from classifier-centric adaptation: the final diagnosis is not produced by a learned classifier head over the query embedding, but by majority voting over the evolved support set,
2
A common misconception is therefore that Evo-RAD is primarily a fine-tuning method; in fact, its defining mechanism is retrieval-time editing of evidence.
3. Graph state representation and policy architecture
Evo-RAD represents the current support set as a graph state
3
so that evidence can be evaluated jointly rather than item by item (Xia et al., 22 Jun 2026). This is important because a reference case may appear plausible relative to the query alone while remaining inconsistent with the support set as a whole.
The graph topology is based on semantic consistency among disease labels. Because raw labels may be too coarse, each label 4 is expanded into fine-grained clinical tags using GPT-5.1,
5
and the tag text is encoded with the frozen text encoder of the vision-LLM to obtain a semantic embedding 6. The adjacency matrix is then defined by cosine similarity,
7
Node features are stored in
8
with each node represented as
9
The statistical descriptor 0 contains four base metrics and their mean deviations, encoding similarity to the query in visual space, the initial retrieval rank, and agreement with the current set in terms of clinical text and disease categories. The inclusion of mean deviations means the policy observes not only absolute relevance but also deviation from the group consensus.
Graph propagation is performed with a two-layer GCN. With self-loops,
1
the update takes the standard normalized form
2
After propagation, action scoring combines node-level local scores for deletion candidates with graph-level global scores for insertion and termination: 3 Pooling is by mean aggregation over nodes.
The architecture therefore encodes a specific inductive bias: evidence quality is a set-level property. DELETE is scored locally because it targets a particular reference, while INSERT and TERMINATE depend on the global coherence of the current set.
4. Learning objective, GRPO optimization, and homogeneity-aware reward
The policy is optimized with Group Relative Policy Optimization (GRPO) rather than an actor-critic scheme (Xia et al., 22 Jun 2026). The paper motivates this choice by noting that rewards are sparse and trajectory quality varies strongly across queries, making critic learning unstable. For each query 4, a group of trajectories
5
is sampled from the old policy, and returns are normalized into relative advantages,
6
The policy objective is a GRPO-style relative policy loss with KL regularization: 7
The reward is explicitly homogeneity-aware. For a trajectory 8, the total return is
9
The trajectory-level reward combines diagnostic correctness, support-set purity, and semantic density. Label purity is
0
and semantic density is
1
The terminal reward is
2
Step-level rewards provide immediate positive feedback for constructive edits: 3 The paper emphasizes a zero-penalty principle: incorrect actions receive reward 4, not a negative penalty. The stated rationale is that negative shaping can hinder early exploration in sparse-reward settings.
A common misunderstanding is that Evo-RAD optimizes retrieval accuracy directly in a static sense. More precisely, it optimizes the homogeneity and diagnostic utility of the final support set under sequential editing. The retrieval target is therefore a coherent evidence cluster rather than merely a similarity-ranked neighbor list.
5. Experimental setup and implementation
Evo-RAD is implemented on top of RetiZero, which remains frozen throughout the method. The trainable component is a small graph policy network with only 116.4K trainable parameters, making the system a lightweight inference-time adapter rather than a full backbone adaptation scheme (Xia et al., 22 Jun 2026).
The paper uses the Retina Image Bank (RIB) rather than heavily used public fundus datasets, on the grounds that strongly exposed benchmarks would obscure genuine generalization. From 30,662 images, a single-label fundus dataset is curated to reduce co-morbidity confounding. Two benchmarks are defined. Rare-20 contains 20 rare diseases, each with 5 images. Retina-31 contains the Rare-20 classes plus 11 common diseases, where common classes satisfy 6 images per class. The train/validation/test split is 70/15/15. For retrieval-based methods, the retrieval corpus is restricted to the training split. The reported split sizes are 524 / 103 / 103 for Rare-20 and 4,737 / 1,001 / 1,023 for Retina-31.
The main baselines fall into three categories. Foundation-model baselines include BiomedCLIP, MedCLIP, EyeCLIP, FLAIR, and RetiZero. PEFT baselines include CoOp, CLIP-Adapter, XCoOp, BiomedCoOp, TDA, Tip-Adapter, and DPC. Retrieval baselines include Static Retrieval and RAC. The core metrics are Accuracy, Macro-F1 score, and Sensitivity, with sensitivity emphasized because the task concerns rare-disease detection under imbalance.
Implementation settings are fixed and compact: candidate buffer size 7, initial retrieval size 8, maximum actions 10, minimum support set size 9, and GRPO group size 0. Training and evaluation use 3-seed averages on an NVIDIA RTX 4090 (24GB). These choices show that the framework is operationally lightweight relative to full fine-tuning, though inference is necessarily more expensive than static retrieval because the support set is edited over multiple steps.
6. Empirical performance, ablations, and limitations
The paper reports that Evo-RAD improves rare-disease diagnosis substantially, with the abstract stating gains of +21.04% over retinal foundation models and +3.56% over retrieval-based and parameter-efficient fine-tuning methods (Xia et al., 22 Jun 2026). In the Rare-20 benchmark, Evo-RAD reaches 46.28% ACC, 40.99% macro-F1, and 42.43% sensitivity. This exceeds RetiZero zero-shot at 25.24 / 26.37 / 29.35, RetiZero linear probe at 26.86 / 10.15 / 13.05, and RAC at 40.06% sensitivity. On Retina-31, Evo-RAD achieves 65.33% ACC, 51.45% macro-F1, and 49.53% sensitivity, surpassing the reported foundation-model, retrieval, and PEFT baselines.
The empirical case for the method rests especially on the comparison between static retrieval and agentic retrieval. The reported gains indicate that simply retrieving similar examples is not enough; the ability to delete discordant evidence and insert additional support produces a more useful diagnostic set. This is most visible in rare-disease sensitivity, where common-disease hubs are especially damaging.
The ablation studies show that the full design is not reducible to a single component. In the state representation, removing mean deviations reduces sensitivity from 42.43% to 36.21%, while removing base metrics reduces it to 38.41%. This supports the claim that both individual relevance and group-level consensus are needed. For the terminal reward, removing 1 yields 36.40% sensitivity, removing 2 yields 33.30%, and removing 3 yields 36.80%. The largest drop comes from removing purity reward, indicating that explicit support-set purification is central rather than peripheral. Step-level rewards are also important: without 4, sensitivity falls to 34.85%, and without 5, to 35.65%. The initial retrieval size is also consequential, with performance peaking at
6
which the authors interpret as a balance between insufficient evidence at small 7 and excessive noise at large 8.
The method has several important limitations. The paper explicitly states that the current framework is designed for single-label diagnosis and that future work will extend it to multi-label and co-morbidity settings. Additional limitations are implied by the design. Evo-RAD depends on the quality of the initial retrieval pool and candidate buffer, relies on labeled training data for RL reward computation, and constructs its semantic graph from GPT-5.1-generated label expansions. It also does not make the inserted candidate itself a learned action; insertion always follows the fixed visual prior of the most similar unseen candidate in the buffer. A plausible implication is that the framework’s flexibility lies primarily in support-set editing policy rather than in full search over candidate identity.
Its practical significance lies in how it redefines retrieval-augmented diagnosis for rare disease. Evo-RAD does not attempt to solve long-tail bias by changing the frozen representation alone. Instead, it treats evidence acquisition as a sequential decision problem, learns to prefer label-pure and semantically dense support sets, and bases prediction on the resulting evidence cluster. Within retinal diagnosis, that makes it a method for retrieval-time adaptation rather than conventional fine-tuning, and a framework whose central object is the evolving support set rather than the classifier head.