---
title: 'RNN-Descent: Fast Graph Construction for ANNS'
url: https://www.emergentmind.com/topics/relative-nearest-neighbor-descent-rnn-descent
type: topic
---

# RNN-Descent: Fast Graph Construction for ANNS

Searching arXiv for the RNN-Descent paper and closely related ANN graph-construction work to ground the article in current literature.
Relative Nearest Neighbor Descent (RNN-Descent) is a direct graph-construction algorithm for Approximate Nearest Neighbor Search (ANNS) that builds a sparse search graph by marrying NN-Descent and the Relative Neighborhood Graph (RNG) strategy [2310.20419]. In graph-based ANNS, which is described as the family of methods with the best balance of accuracy and speed for million-scale datasets, the principal disadvantage is long index construction time; RNN-Descent targets that bottleneck by directly constructing a graph-based index without first building an approximate $K$-nearest neighbor graph and without costly internal ANNS calls [2310.20419]. Experimental results reported for GIST1M indicate that its construction is $2\times$ faster than NSG, while search performance is comparable to existing state-of-the-art methods such as NSG, and it is even faster than the construction speed of NN-Descent [2310.20419].

## 1. Problem formulation and design objective

Let $X=\{x_1,\dots,x_n\}\subset\mathbb R^d$ be the database, with metric $\delta(u,v)=\|x_u-x_v\|$, for example Euclidean distance. ANNS aims to find, for a query $q$, $\arg\min_i \delta(q,x_i)$ [2310.20419]. Within this setting, RNN-Descent is designed as a sparse graph-construction method for graph-based ANNS rather than as a general-purpose exact neighborhood graph algorithm.

The central objective is to accelerate index construction. The motivating observation is that many methods had improved the tradeoff between accuracy and speed during search, whereas there had been comparatively little research on accelerating construction [2310.20419]. RNN-Descent therefore emphasizes direct construction of the search graph itself. Unlike refinement-based methods such as NSG that first build an approximate $K$-NN graph and then prune it, RNN-Descent initializes a random sparse graph and simultaneously adds useful edges and prunes redundant edges [2310.20419].

This suggests a shift in emphasis from two-stage graph generation toward a single iterative procedure in which edge discovery and edge selection are coupled. A plausible implication is that the main efficiency gain derives not from changing graph traversal at query time, but from reducing construction stages and avoiding an expensive full intermediate $K$-NN structure.

## 2. Algorithmic antecedents: NN-Descent and RNG pruning

RNN-Descent combines two ingredients. The first is NN-Descent, described in the source material as an efficient local search for $K$-nearest neighbors via iteratively “joining” neighbor lists [2310.20419]. In NN-Descent, each node $u$ maintains a candidate neighbor set $N(u)$ of size $K$. The key insight is that if $v\in N(u)$ and $w\in N(u)$, then $(v,w)$ is a good candidate pair. In each iteration, the algorithm performs a “join” step over pairs in $N(u)\times N(u)$ involving at least one newly added neighbor, evaluates $\delta(v,w)$, and inserts $(v,w)$ and $(w,v)$ into candidate lists, followed by an “update” step that keeps the top $K$ by distance [2310.20419].

The second ingredient is the RNG strategy. The Relative Neighborhood Graph of $X$ contains edge $(u,v)$ only if there is no third point $w$ such that
$$
\max\{\delta(u,w),\delta(v,w)\}<\delta(u,v).
$$
Equivalently, during local pruning of a candidate set $U$ for a node $u$, one can sort $U$ by increasing $\delta(u,\cdot)$ and keep $v\in U$ only if for every already selected neighbor $w$,
$$
\delta(u,v)<\delta(v,w).
$$
This criterion removes “long” edges that are redundant for greedy search [2310.20419].

In the later GPU-oriented formulation, the same pruning principle is expressed as: for central vertex $v$ and candidates $n_1,n_2\in N_v$, both edges are retained if and only if
$$
d(v,n_1)<d(n_1,n_2)\quad\land\quad d(v,n_2)<d(n_1,n_2).
$$
If instead $d(n_1,n_2)<d(v,n_2)$, then $n_2$ is redirected to $n_1$ [2510.02774].

Taken together, these components define the conceptual identity of RNN-Descent: NN-Descent supplies neighbor-of-neighbor proposals, while RNG supplies a geometric pruning rule that preferentially preserves edges effective for greedy graph search [2310.20419].

## 3. Core procedure of RNN-Descent

RNN-Descent alternates two subroutines: **UpdateNeighbors**, which integrates NN-Descent proposals with RNG pruning, and **AddReverseEdges**, which injects reverse edges to avoid getting stuck in an RNG-local optimum [2310.20419]. The overall procedure uses parameters $S$ for initial degree, $R$ for maximum degree after pruning, and $(T_1,T_2)$ for the number of outer and inner iterations.

Initialization constructs a random directed $S$-regular graph and marks all edges as “new.” For each outer round $t_1=1,\dots,T_1$, the algorithm runs UpdateNeighbors for $T_2$ inner passes and, if $t_1<T_1$, performs AddReverseEdges. The final graph $G$ is then returned [2310.20419].

In **UpdateNeighbors**, for each node $u$, the outgoing neighbors are sorted by increasing $\delta(u,v)$. The algorithm scans candidates in that order, maintaining a selected set $U'$. For a candidate $v$, it compares $v$ against already selected neighbors $w\in U'$. If both $v$ and $w$ are “old,” the comparison is skipped because it has already been done. Otherwise, if $\delta(u,v)\ge \delta(v,w)$, then $(u,v)$ fails the RNG test: the edge $(u,v)$ is removed, edge $(w,v)$ is inserted if absent, and the scan for that candidate stops. If no such blocking neighbor is found, $v$ is added to $U'$. After processing, all edges from $u$ to nodes in $U'$ are marked “old” [2310.20419].

This single step performs two tasks simultaneously. It implements NN-Descent style edge proposals because when $(u,v)$ fails the RNG test against $w$, it proposes $(w,v)$. It also performs RNG-style pruning because $(u,v)$ is removed whenever $\delta(v,w)\le \delta(u,v)$ [2310.20419].

In **AddReverseEdges**, every existing edge $(u,v)$ is reversed to add $(v,u)$, with newly added edges marked “new.” Then, for each node $u$, incoming neighbors are sorted by $\delta(v,u)$ and only the $R$ smallest are kept, enforcing in-degree $\le R$; the out-degree is similarly pruned to enforce out-degree $\le R$ [2310.20419]. The stated purpose of this operation is to escape local optima and preserve connectivity.

At query time, search on the constructed graph uses standard graph-traversal ANNS, but at each visited node $u$ it expands only the top $K$ neighbors, by $\delta(u,\cdot)$, among the outgoing edges [2310.20419].

## 4. Complexity, parameterization, and reported empirical behavior

Let $n=|V|$ and let $D\approx$ average degree $\approx R$. Each UpdateNeighbors pass visits every node and its $D$ neighbors, and for each candidate compares against up to $D'\lesssim D$ selected neighbors, giving $O(n\cdot D^2)$ per pass. With $T_2$ inner passes and $T_1$ outer rounds, and with AddReverseEdges costing $O(n\cdot D\cdot \log D)$ each time it occurs, the overall cost is approximately
$$
O(n\cdot D^2\cdot T_2\cdot T_1 + n\cdot D\cdot \log D\cdot T_1)
$$
[2310.20419].

The source contrasts this with standard NN-Descent and NSG. Standard NN-Descent with $K=64$ and $10$ iterations costs roughly $O(n\cdot 64^2\cdot 10)=O(40960n)$ distance evaluations. NSG first runs NN-Descent to build the $K$-NN graph and then applies RNG-style refinement $O(nK^2)$. HNSW inserts $n$ points by ANNS of cost $O(\log n)$ each, giving $O(n\log n\cdot c)$ with a large constant $c$ due to search accuracy demands [2310.20419].

The practical regime emphasized for RNN-Descent is one where $D\ll K$ of a full $K$-NN graph and $T_1\cdot T_2$ is modest, with the example $(T_1,T_2)=(4,15)$ highlighted as striking a good balance [2310.20419]. Parameter guidance in the same source specifies that a small initial degree $S$ in the range $10$–$30$ is sufficient to bootstrap connectivity; $R$ controls sparsity and is typically about $3S$–$5S$; larger $T_2$ accelerates local convergence; larger $T_1$ allows more round resets for global quality; and search-time expansion parameter $K$ is a post-construction speed–recall knob, with examples $K=16\ldots 64$ [2310.20419].

For GIST1M, with $n=1$M and $d=960$, the reported search curves of RNN-Descent are virtually identical to NSG and HNSW for recall up to $0.95$. Construction times are given as approximately $2\,000$ s for NSG, approximately $1\,400$ s for NN-Descent, and approximately $1\,000$ s for RNN-Descent. The graph’s average out-degree is reported as approximately $20$, stated to be the same order as NSG’s approximately $10$–$20$, ensuring low memory cost [2310.20419].

These results support the characterization of RNN-Descent as a direct-construction ANNS graph that requires no prior $K$-NN graph, needs no costly internal ANNS calls, converges in $O(n\cdot D^2\cdot T)$ time with small $D$ and modest $T$, and matches or exceeds the search performance of state-of-the-art graph indexes while cutting index build time by up to half [2310.20419].

## 5. GPU parallelization: GRNND

The later paper "GRNND: A GPU-Parallel Relative NN-Descent Algorithm for Efficient Approximate Nearest Neighbor Graph Construction" presents the first GPU-parallel algorithm of RNN-Descent designed to fully exploit GPU architecture [2510.02774]. Its starting point is the observation that as data amount and dimensionality increase, the complexity of graph construction in RNN-Descent rises sharply, making this stage increasingly time-consuming and even prohibitive for subsequent query processing.

The GRNND design introduces four mechanisms. First, **disordered neighbor propagation** replaces strict ascending-distance processing with random neighbor-pair sampling from $N_v$, followed by RNG tests in arbitrary order; this is stated to mitigate synchronized update traps, enhance structural diversity, and avoid premature convergence during parallel execution [2510.02774]. Second, **warp-level cooperative execution** assigns one CUDA warp of $32$ threads to each vertex update; distance computation uses striped loading over dimensions with warp-wide reduction via `__shfl_down`, while duplicate detection and insertion use `__ballot` and warp-max reductions rather than atomics or global locks [2510.02774]. Third, **fixed-capacity double-buffered pools** provide each vertex with two static arrays of length $R$, one read set and one write set, which are swapped after each inner iteration; this removes dynamic allocation and lock contention [2510.02774]. Fourth, **reverse-edge sampling** adds only a fraction $\rho$ of reverse edges, rather than all of them, to control memory growth and contention while maintaining structural diversity [2510.02774].

In the GRNND formulation, one update has complexity $O(R\log R + R^2\cdot d)\approx O(R^2d)$, so over $n$ vertices and $T_1\cdot T_2$ total refinement passes the total time is
$$
T_{\rm total}=O(T_1\cdot T_2\cdot n\cdot R^2\cdot d),
$$
with space complexity $O(nR)$ [2510.02774]. The same source notes empirically that small $T_1,T_2$ can suffice for high recall, citing the example $T_1=3,\,T_2=4$ [2510.02774].

Experiments compare GRNND with GPU baselines GANNS, CAGRA, and GGNN, and CPU baselines RNN-Descent(CPU) and HNSW on SIFT1M, DEEP1M, and GIST1M, using RTX 4090 and RTX 6000 Ada GPUs [2510.02774]. At Recall@10 near $0.90$ on RTX 4090, construction times are reported as follows: on SIFT1M, GANNS $1.56$ s, CAGRA $4.42$ s, GGNN $20.2$ s, RNN-Descent (CPU) $7.16$ s, HNSW $24.5$ s, and GRNND $0.58$ s; on DEEP1M, the corresponding times are $1.66$ s, $4.82$ s, $18.8$ s, $21.6$ s, $43.9$ s, and $0.73$ s; on GIST1M, they are $7.46$ s, $208$ s, $786$ s, $251$ s, $153$ s, and $4.88$ s [2510.02774]. The reported speedups are $2.4\times$–$51.7\times$ over existing GPU methods and $17.8\times$–$49.8\times$ over CPU methods [2510.02774].

Using a common CPU search routine over GPU-built indices, GRNND’s graphs are reported to achieve nearly identical Recall–QPS trade-offs to CPU RNN-Descent, while outperforming HNSW, GGNN, and CAGRA [2510.02774]. The ablations further state that ascending update order suffers convergence traps in parallel, descending order improves recall but at approximately $2\times$ time cost, and disordered propagation gives the best balance, reaching recall near $0.90$ with only approximately $1.2\times$ the time of ascending order. For reverse-edge sampling, recall-versus-time curves on SIFT1M show an optimum at $\rho\approx 0.6$; low $\rho=0.2$ builds fastest but yields recall around $0.85$, while high $\rho=1.0$ marginally raises recall to around $0.95$ but costs approximately $30\%$ more time. For iteration counts, the reported recommendation is $T_1=1, T_2=6$ on SIFT1M, whereas on GIST1M values $T_1=3$–$4$, $T_2=8$ improve recall at moderate cost [2510.02774].

## 6. Position within graph-based ANNS, distinctions, and open questions

RNN-Descent should be distinguished from two-stage refinement pipelines. The source explicitly describes it as a direct-construction algorithm that combines NN-Descent and RNG strategy so that useful edges are added and redundant edges pruned in the same iterative process, rather than by first materializing an approximate $K$-NN graph and then refining it [2310.20419]. That distinction is central to its reported construction-time advantage.

It should also be distinguished from the view that graph quality must be traded away to obtain faster construction. On GIST1M, the original paper reports search curves virtually identical to NSG and HNSW up to recall $0.95$ while reducing construction time relative to both NSG and NN-Descent [2310.20419]. The GPU study makes the same point in a different computational regime: GRNND is reported to deliver large construction-time speedups without loss of Search-Recall quality relative to CPU RNN-Descent [2510.02774].

At the same time, the later work identifies explicit limitations and open directions. Fixed $R$ may limit maximum recall in ultra-sparse regimes; multi-GPU or distributed extensions are described as non-trivial because of data-exchange patterns; and integration into billion-scale, out-of-core pipelines such as GPU clusters remains future work [2510.02774]. These statements delimit the current scope of the method family: the construction-time improvements are well established for the reported sparse-graph setting, but extension to broader hardware and larger-scale deployment remains an open research problem.

In this sense, RNN-Descent occupies a specific niche within graph-based ANNS: it is a sparse, direct-construction graph builder whose defining mechanism is the fusion of neighbor-of-neighbor proposal generation with RNG-based pruning, and whose later GPU realization preserves that structure while replacing sequential control flow and dynamic memory with randomized update order, warp-cooperative execution, double-buffered fixed-capacity pools, and sampled reverse-edge insertion [2310.20419].

Source: https://www.emergentmind.com/topics/relative-nearest-neighbor-descent-rnn-descent