Papers
Topics
Authors
Recent
Search
2000 character limit reached

RNN-Descent: Fast Graph Construction for ANNS

Updated 14 July 2026
  • RNN-Descent is a direct graph-construction algorithm for ANNS that fuses NN-Descent proposals with RNG pruning to build sparse yet effective search graphs.
  • It accelerates index construction by eliminating the costly intermediate K-NN graph phase, achieving up to 2x faster build times on datasets like GIST1M while maintaining high search accuracy.
  • Recent GPU adaptations such as GRNND leverage parallelism to deliver significant speedups (up to 51.7x) and maintain competitive recall, demonstrating scalability for high-dimensional data.

Searching arXiv for the RNN-Descent paper and closely related ANN graph-construction work to ground the article in current literature. Relative Nearest Neighbor Descent (RNN-Descent) is a direct graph-construction algorithm for Approximate Nearest Neighbor Search (ANNS) that builds a sparse search graph by marrying NN-Descent and the Relative Neighborhood Graph (RNG) strategy (Ono et al., 2023). In graph-based ANNS, which is described as the family of methods with the best balance of accuracy and speed for million-scale datasets, the principal disadvantage is long index construction time; RNN-Descent targets that bottleneck by directly constructing a graph-based index without first building an approximate KK-nearest neighbor graph and without costly internal ANNS calls (Ono et al., 2023). Experimental results reported for GIST1M indicate that its construction is 2×2\times faster than NSG, while search performance is comparable to existing state-of-the-art methods such as NSG, and it is even faster than the construction speed of NN-Descent (Ono et al., 2023).

1. Problem formulation and design objective

Let X={x1,…,xn}⊂RdX=\{x_1,\dots,x_n\}\subset\mathbb R^d be the database, with metric δ(u,v)=∥xu−xv∥\delta(u,v)=\|x_u-x_v\|, for example Euclidean distance. ANNS aims to find, for a query qq, arg⁡min⁡iδ(q,xi)\arg\min_i \delta(q,x_i) (Ono et al., 2023). Within this setting, RNN-Descent is designed as a sparse graph-construction method for graph-based ANNS rather than as a general-purpose exact neighborhood graph algorithm.

The central objective is to accelerate index construction. The motivating observation is that many methods had improved the tradeoff between accuracy and speed during search, whereas there had been comparatively little research on accelerating construction (Ono et al., 2023). RNN-Descent therefore emphasizes direct construction of the search graph itself. Unlike refinement-based methods such as NSG that first build an approximate KK-NN graph and then prune it, RNN-Descent initializes a random sparse graph and simultaneously adds useful edges and prunes redundant edges (Ono et al., 2023).

This suggests a shift in emphasis from two-stage graph generation toward a single iterative procedure in which edge discovery and edge selection are coupled. A plausible implication is that the main efficiency gain derives not from changing graph traversal at query time, but from reducing construction stages and avoiding an expensive full intermediate KK-NN structure.

2. Algorithmic antecedents: NN-Descent and RNG pruning

RNN-Descent combines two ingredients. The first is NN-Descent, described in the source material as an efficient local search for KK-nearest neighbors via iteratively “joining” neighbor lists (Ono et al., 2023). In NN-Descent, each node uu maintains a candidate neighbor set 2×2\times0 of size 2×2\times1. The key insight is that if 2×2\times2 and 2×2\times3, then 2×2\times4 is a good candidate pair. In each iteration, the algorithm performs a “join” step over pairs in 2×2\times5 involving at least one newly added neighbor, evaluates 2×2\times6, and inserts 2×2\times7 and 2×2\times8 into candidate lists, followed by an “update” step that keeps the top 2×2\times9 by distance (Ono et al., 2023).

The second ingredient is the RNG strategy. The Relative Neighborhood Graph of X={x1,…,xn}⊂RdX=\{x_1,\dots,x_n\}\subset\mathbb R^d0 contains edge X={x1,…,xn}⊂RdX=\{x_1,\dots,x_n\}\subset\mathbb R^d1 only if there is no third point X={x1,…,xn}⊂RdX=\{x_1,\dots,x_n\}\subset\mathbb R^d2 such that

X={x1,…,xn}⊂RdX=\{x_1,\dots,x_n\}\subset\mathbb R^d3

Equivalently, during local pruning of a candidate set X={x1,…,xn}⊂RdX=\{x_1,\dots,x_n\}\subset\mathbb R^d4 for a node X={x1,…,xn}⊂RdX=\{x_1,\dots,x_n\}\subset\mathbb R^d5, one can sort X={x1,…,xn}⊂RdX=\{x_1,\dots,x_n\}\subset\mathbb R^d6 by increasing X={x1,…,xn}⊂RdX=\{x_1,\dots,x_n\}\subset\mathbb R^d7 and keep X={x1,…,xn}⊂RdX=\{x_1,\dots,x_n\}\subset\mathbb R^d8 only if for every already selected neighbor X={x1,…,xn}⊂RdX=\{x_1,\dots,x_n\}\subset\mathbb R^d9,

δ(u,v)=∥xu−xv∥\delta(u,v)=\|x_u-x_v\|0

This criterion removes “long” edges that are redundant for greedy search (Ono et al., 2023).

In the later GPU-oriented formulation, the same pruning principle is expressed as: for central vertex δ(u,v)=∥xu−xv∥\delta(u,v)=\|x_u-x_v\|1 and candidates δ(u,v)=∥xu−xv∥\delta(u,v)=\|x_u-x_v\|2, both edges are retained if and only if

δ(u,v)=∥xu−xv∥\delta(u,v)=\|x_u-x_v\|3

If instead δ(u,v)=∥xu−xv∥\delta(u,v)=\|x_u-x_v\|4, then δ(u,v)=∥xu−xv∥\delta(u,v)=\|x_u-x_v\|5 is redirected to δ(u,v)=∥xu−xv∥\delta(u,v)=\|x_u-x_v\|6 (Li et al., 3 Oct 2025).

Taken together, these components define the conceptual identity of RNN-Descent: NN-Descent supplies neighbor-of-neighbor proposals, while RNG supplies a geometric pruning rule that preferentially preserves edges effective for greedy graph search (Ono et al., 2023).

3. Core procedure of RNN-Descent

RNN-Descent alternates two subroutines: UpdateNeighbors, which integrates NN-Descent proposals with RNG pruning, and AddReverseEdges, which injects reverse edges to avoid getting stuck in an RNG-local optimum (Ono et al., 2023). The overall procedure uses parameters δ(u,v)=∥xu−xv∥\delta(u,v)=\|x_u-x_v\|7 for initial degree, δ(u,v)=∥xu−xv∥\delta(u,v)=\|x_u-x_v\|8 for maximum degree after pruning, and δ(u,v)=∥xu−xv∥\delta(u,v)=\|x_u-x_v\|9 for the number of outer and inner iterations.

Initialization constructs a random directed qq0-regular graph and marks all edges as “new.” For each outer round qq1, the algorithm runs UpdateNeighbors for qq2 inner passes and, if qq3, performs AddReverseEdges. The final graph qq4 is then returned (Ono et al., 2023).

In UpdateNeighbors, for each node qq5, the outgoing neighbors are sorted by increasing qq6. The algorithm scans candidates in that order, maintaining a selected set qq7. For a candidate qq8, it compares qq9 against already selected neighbors arg⁡min⁡iδ(q,xi)\arg\min_i \delta(q,x_i)0. If both arg⁡min⁡iδ(q,xi)\arg\min_i \delta(q,x_i)1 and arg⁡min⁡iδ(q,xi)\arg\min_i \delta(q,x_i)2 are “old,” the comparison is skipped because it has already been done. Otherwise, if arg⁡min⁡iδ(q,xi)\arg\min_i \delta(q,x_i)3, then arg⁡min⁡iδ(q,xi)\arg\min_i \delta(q,x_i)4 fails the RNG test: the edge arg⁡min⁡iδ(q,xi)\arg\min_i \delta(q,x_i)5 is removed, edge arg⁡min⁡iδ(q,xi)\arg\min_i \delta(q,x_i)6 is inserted if absent, and the scan for that candidate stops. If no such blocking neighbor is found, arg⁡min⁡iδ(q,xi)\arg\min_i \delta(q,x_i)7 is added to arg⁡min⁡iδ(q,xi)\arg\min_i \delta(q,x_i)8. After processing, all edges from arg⁡min⁡iδ(q,xi)\arg\min_i \delta(q,x_i)9 to nodes in KK0 are marked “old” (Ono et al., 2023).

This single step performs two tasks simultaneously. It implements NN-Descent style edge proposals because when KK1 fails the RNG test against KK2, it proposes KK3. It also performs RNG-style pruning because KK4 is removed whenever KK5 (Ono et al., 2023).

In AddReverseEdges, every existing edge KK6 is reversed to add KK7, with newly added edges marked “new.” Then, for each node KK8, incoming neighbors are sorted by KK9 and only the KK0 smallest are kept, enforcing in-degree KK1; the out-degree is similarly pruned to enforce out-degree KK2 (Ono et al., 2023). The stated purpose of this operation is to escape local optima and preserve connectivity.

At query time, search on the constructed graph uses standard graph-traversal ANNS, but at each visited node KK3 it expands only the top KK4 neighbors, by KK5, among the outgoing edges (Ono et al., 2023).

4. Complexity, parameterization, and reported empirical behavior

Let KK6 and let KK7 average degree KK8. Each UpdateNeighbors pass visits every node and its KK9 neighbors, and for each candidate compares against up to KK0 selected neighbors, giving KK1 per pass. With KK2 inner passes and KK3 outer rounds, and with AddReverseEdges costing KK4 each time it occurs, the overall cost is approximately

KK5

(Ono et al., 2023).

The source contrasts this with standard NN-Descent and NSG. Standard NN-Descent with KK6 and KK7 iterations costs roughly KK8 distance evaluations. NSG first runs NN-Descent to build the KK9-NN graph and then applies RNG-style refinement uu0. HNSW inserts uu1 points by ANNS of cost uu2 each, giving uu3 with a large constant uu4 due to search accuracy demands (Ono et al., 2023).

The practical regime emphasized for RNN-Descent is one where uu5 of a full uu6-NN graph and uu7 is modest, with the example uu8 highlighted as striking a good balance (Ono et al., 2023). Parameter guidance in the same source specifies that a small initial degree uu9 in the range 2×2\times00–2×2\times01 is sufficient to bootstrap connectivity; 2×2\times02 controls sparsity and is typically about 2×2\times03–2×2\times04; larger 2×2\times05 accelerates local convergence; larger 2×2\times06 allows more round resets for global quality; and search-time expansion parameter 2×2\times07 is a post-construction speed–recall knob, with examples 2×2\times08 (Ono et al., 2023).

For GIST1M, with 2×2\times09M and 2×2\times10, the reported search curves of RNN-Descent are virtually identical to NSG and HNSW for recall up to 2×2\times11. Construction times are given as approximately 2×2\times12 s for NSG, approximately 2×2\times13 s for NN-Descent, and approximately 2×2\times14 s for RNN-Descent. The graph’s average out-degree is reported as approximately 2×2\times15, stated to be the same order as NSG’s approximately 2×2\times16–2×2\times17, ensuring low memory cost (Ono et al., 2023).

These results support the characterization of RNN-Descent as a direct-construction ANNS graph that requires no prior 2×2\times18-NN graph, needs no costly internal ANNS calls, converges in 2×2\times19 time with small 2×2\times20 and modest 2×2\times21, and matches or exceeds the search performance of state-of-the-art graph indexes while cutting index build time by up to half (Ono et al., 2023).

5. GPU parallelization: GRNND

The later paper "GRNND: A GPU-Parallel Relative NN-Descent Algorithm for Efficient Approximate Nearest Neighbor Graph Construction" presents the first GPU-parallel algorithm of RNN-Descent designed to fully exploit GPU architecture (Li et al., 3 Oct 2025). Its starting point is the observation that as data amount and dimensionality increase, the complexity of graph construction in RNN-Descent rises sharply, making this stage increasingly time-consuming and even prohibitive for subsequent query processing.

The GRNND design introduces four mechanisms. First, disordered neighbor propagation replaces strict ascending-distance processing with random neighbor-pair sampling from 2×2\times22, followed by RNG tests in arbitrary order; this is stated to mitigate synchronized update traps, enhance structural diversity, and avoid premature convergence during parallel execution (Li et al., 3 Oct 2025). Second, warp-level cooperative execution assigns one CUDA warp of 2×2\times23 threads to each vertex update; distance computation uses striped loading over dimensions with warp-wide reduction via __shfl_down, while duplicate detection and insertion use __ballot and warp-max reductions rather than atomics or global locks (Li et al., 3 Oct 2025). Third, fixed-capacity double-buffered pools provide each vertex with two static arrays of length 2×2\times24, one read set and one write set, which are swapped after each inner iteration; this removes dynamic allocation and lock contention (Li et al., 3 Oct 2025). Fourth, reverse-edge sampling adds only a fraction 2×2\times25 of reverse edges, rather than all of them, to control memory growth and contention while maintaining structural diversity (Li et al., 3 Oct 2025).

In the GRNND formulation, one update has complexity 2×2\times26, so over 2×2\times27 vertices and 2×2\times28 total refinement passes the total time is

2×2\times29

with space complexity 2×2\times30 (Li et al., 3 Oct 2025). The same source notes empirically that small 2×2\times31 can suffice for high recall, citing the example 2×2\times32 (Li et al., 3 Oct 2025).

Experiments compare GRNND with GPU baselines GANNS, CAGRA, and GGNN, and CPU baselines RNN-Descent(CPU) and HNSW on SIFT1M, DEEP1M, and GIST1M, using RTX 4090 and RTX 6000 Ada GPUs (Li et al., 3 Oct 2025). At Recall@10 near 2×2\times33 on RTX 4090, construction times are reported as follows: on SIFT1M, GANNS 2×2\times34 s, CAGRA 2×2\times35 s, GGNN 2×2\times36 s, RNN-Descent (CPU) 2×2\times37 s, HNSW 2×2\times38 s, and GRNND 2×2\times39 s; on DEEP1M, the corresponding times are 2×2\times40 s, 2×2\times41 s, 2×2\times42 s, 2×2\times43 s, 2×2\times44 s, and 2×2\times45 s; on GIST1M, they are 2×2\times46 s, 2×2\times47 s, 2×2\times48 s, 2×2\times49 s, 2×2\times50 s, and 2×2\times51 s (Li et al., 3 Oct 2025). The reported speedups are 2×2\times52–2×2\times53 over existing GPU methods and 2×2\times54–2×2\times55 over CPU methods (Li et al., 3 Oct 2025).

Using a common CPU search routine over GPU-built indices, GRNND’s graphs are reported to achieve nearly identical Recall–QPS trade-offs to CPU RNN-Descent, while outperforming HNSW, GGNN, and CAGRA (Li et al., 3 Oct 2025). The ablations further state that ascending update order suffers convergence traps in parallel, descending order improves recall but at approximately 2×2\times56 time cost, and disordered propagation gives the best balance, reaching recall near 2×2\times57 with only approximately 2×2\times58 the time of ascending order. For reverse-edge sampling, recall-versus-time curves on SIFT1M show an optimum at 2×2\times59; low 2×2\times60 builds fastest but yields recall around 2×2\times61, while high 2×2\times62 marginally raises recall to around 2×2\times63 but costs approximately 2×2\times64 more time. For iteration counts, the reported recommendation is 2×2\times65 on SIFT1M, whereas on GIST1M values 2×2\times66–2×2\times67, 2×2\times68 improve recall at moderate cost (Li et al., 3 Oct 2025).

6. Position within graph-based ANNS, distinctions, and open questions

RNN-Descent should be distinguished from two-stage refinement pipelines. The source explicitly describes it as a direct-construction algorithm that combines NN-Descent and RNG strategy so that useful edges are added and redundant edges pruned in the same iterative process, rather than by first materializing an approximate 2×2\times69-NN graph and then refining it (Ono et al., 2023). That distinction is central to its reported construction-time advantage.

It should also be distinguished from the view that graph quality must be traded away to obtain faster construction. On GIST1M, the original paper reports search curves virtually identical to NSG and HNSW up to recall 2×2\times70 while reducing construction time relative to both NSG and NN-Descent (Ono et al., 2023). The GPU study makes the same point in a different computational regime: GRNND is reported to deliver large construction-time speedups without loss of Search-Recall quality relative to CPU RNN-Descent (Li et al., 3 Oct 2025).

At the same time, the later work identifies explicit limitations and open directions. Fixed 2×2\times71 may limit maximum recall in ultra-sparse regimes; multi-GPU or distributed extensions are described as non-trivial because of data-exchange patterns; and integration into billion-scale, out-of-core pipelines such as GPU clusters remains future work (Li et al., 3 Oct 2025). These statements delimit the current scope of the method family: the construction-time improvements are well established for the reported sparse-graph setting, but extension to broader hardware and larger-scale deployment remains an open research problem.

In this sense, RNN-Descent occupies a specific niche within graph-based ANNS: it is a sparse, direct-construction graph builder whose defining mechanism is the fusion of neighbor-of-neighbor proposal generation with RNG-based pruning, and whose later GPU realization preserves that structure while replacing sequential control flow and dynamic memory with randomized update order, warp-cooperative execution, double-buffered fixed-capacity pools, and sampled reverse-edge insertion (Ono et al., 2023).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Relative Nearest Neighbor Descent (RNN-Descent).