RNN-Descent: Fast Graph Construction for ANNS
- RNN-Descent is a direct graph-construction algorithm for ANNS that fuses NN-Descent proposals with RNG pruning to build sparse yet effective search graphs.
- It accelerates index construction by eliminating the costly intermediate K-NN graph phase, achieving up to 2x faster build times on datasets like GIST1M while maintaining high search accuracy.
- Recent GPU adaptations such as GRNND leverage parallelism to deliver significant speedups (up to 51.7x) and maintain competitive recall, demonstrating scalability for high-dimensional data.
Searching arXiv for the RNN-Descent paper and closely related ANN graph-construction work to ground the article in current literature. Relative Nearest Neighbor Descent (RNN-Descent) is a direct graph-construction algorithm for Approximate Nearest Neighbor Search (ANNS) that builds a sparse search graph by marrying NN-Descent and the Relative Neighborhood Graph (RNG) strategy (Ono et al., 2023). In graph-based ANNS, which is described as the family of methods with the best balance of accuracy and speed for million-scale datasets, the principal disadvantage is long index construction time; RNN-Descent targets that bottleneck by directly constructing a graph-based index without first building an approximate -nearest neighbor graph and without costly internal ANNS calls (Ono et al., 2023). Experimental results reported for GIST1M indicate that its construction is faster than NSG, while search performance is comparable to existing state-of-the-art methods such as NSG, and it is even faster than the construction speed of NN-Descent (Ono et al., 2023).
1. Problem formulation and design objective
Let be the database, with metric , for example Euclidean distance. ANNS aims to find, for a query , (Ono et al., 2023). Within this setting, RNN-Descent is designed as a sparse graph-construction method for graph-based ANNS rather than as a general-purpose exact neighborhood graph algorithm.
The central objective is to accelerate index construction. The motivating observation is that many methods had improved the tradeoff between accuracy and speed during search, whereas there had been comparatively little research on accelerating construction (Ono et al., 2023). RNN-Descent therefore emphasizes direct construction of the search graph itself. Unlike refinement-based methods such as NSG that first build an approximate -NN graph and then prune it, RNN-Descent initializes a random sparse graph and simultaneously adds useful edges and prunes redundant edges (Ono et al., 2023).
This suggests a shift in emphasis from two-stage graph generation toward a single iterative procedure in which edge discovery and edge selection are coupled. A plausible implication is that the main efficiency gain derives not from changing graph traversal at query time, but from reducing construction stages and avoiding an expensive full intermediate -NN structure.
2. Algorithmic antecedents: NN-Descent and RNG pruning
RNN-Descent combines two ingredients. The first is NN-Descent, described in the source material as an efficient local search for -nearest neighbors via iteratively “joining” neighbor lists (Ono et al., 2023). In NN-Descent, each node maintains a candidate neighbor set 0 of size 1. The key insight is that if 2 and 3, then 4 is a good candidate pair. In each iteration, the algorithm performs a “join” step over pairs in 5 involving at least one newly added neighbor, evaluates 6, and inserts 7 and 8 into candidate lists, followed by an “update” step that keeps the top 9 by distance (Ono et al., 2023).
The second ingredient is the RNG strategy. The Relative Neighborhood Graph of 0 contains edge 1 only if there is no third point 2 such that
3
Equivalently, during local pruning of a candidate set 4 for a node 5, one can sort 6 by increasing 7 and keep 8 only if for every already selected neighbor 9,
0
This criterion removes “long” edges that are redundant for greedy search (Ono et al., 2023).
In the later GPU-oriented formulation, the same pruning principle is expressed as: for central vertex 1 and candidates 2, both edges are retained if and only if
3
If instead 4, then 5 is redirected to 6 (Li et al., 3 Oct 2025).
Taken together, these components define the conceptual identity of RNN-Descent: NN-Descent supplies neighbor-of-neighbor proposals, while RNG supplies a geometric pruning rule that preferentially preserves edges effective for greedy graph search (Ono et al., 2023).
3. Core procedure of RNN-Descent
RNN-Descent alternates two subroutines: UpdateNeighbors, which integrates NN-Descent proposals with RNG pruning, and AddReverseEdges, which injects reverse edges to avoid getting stuck in an RNG-local optimum (Ono et al., 2023). The overall procedure uses parameters 7 for initial degree, 8 for maximum degree after pruning, and 9 for the number of outer and inner iterations.
Initialization constructs a random directed 0-regular graph and marks all edges as “new.” For each outer round 1, the algorithm runs UpdateNeighbors for 2 inner passes and, if 3, performs AddReverseEdges. The final graph 4 is then returned (Ono et al., 2023).
In UpdateNeighbors, for each node 5, the outgoing neighbors are sorted by increasing 6. The algorithm scans candidates in that order, maintaining a selected set 7. For a candidate 8, it compares 9 against already selected neighbors 0. If both 1 and 2 are “old,” the comparison is skipped because it has already been done. Otherwise, if 3, then 4 fails the RNG test: the edge 5 is removed, edge 6 is inserted if absent, and the scan for that candidate stops. If no such blocking neighbor is found, 7 is added to 8. After processing, all edges from 9 to nodes in 0 are marked “old” (Ono et al., 2023).
This single step performs two tasks simultaneously. It implements NN-Descent style edge proposals because when 1 fails the RNG test against 2, it proposes 3. It also performs RNG-style pruning because 4 is removed whenever 5 (Ono et al., 2023).
In AddReverseEdges, every existing edge 6 is reversed to add 7, with newly added edges marked “new.” Then, for each node 8, incoming neighbors are sorted by 9 and only the 0 smallest are kept, enforcing in-degree 1; the out-degree is similarly pruned to enforce out-degree 2 (Ono et al., 2023). The stated purpose of this operation is to escape local optima and preserve connectivity.
At query time, search on the constructed graph uses standard graph-traversal ANNS, but at each visited node 3 it expands only the top 4 neighbors, by 5, among the outgoing edges (Ono et al., 2023).
4. Complexity, parameterization, and reported empirical behavior
Let 6 and let 7 average degree 8. Each UpdateNeighbors pass visits every node and its 9 neighbors, and for each candidate compares against up to 0 selected neighbors, giving 1 per pass. With 2 inner passes and 3 outer rounds, and with AddReverseEdges costing 4 each time it occurs, the overall cost is approximately
5
The source contrasts this with standard NN-Descent and NSG. Standard NN-Descent with 6 and 7 iterations costs roughly 8 distance evaluations. NSG first runs NN-Descent to build the 9-NN graph and then applies RNG-style refinement 0. HNSW inserts 1 points by ANNS of cost 2 each, giving 3 with a large constant 4 due to search accuracy demands (Ono et al., 2023).
The practical regime emphasized for RNN-Descent is one where 5 of a full 6-NN graph and 7 is modest, with the example 8 highlighted as striking a good balance (Ono et al., 2023). Parameter guidance in the same source specifies that a small initial degree 9 in the range 00–01 is sufficient to bootstrap connectivity; 02 controls sparsity and is typically about 03–04; larger 05 accelerates local convergence; larger 06 allows more round resets for global quality; and search-time expansion parameter 07 is a post-construction speed–recall knob, with examples 08 (Ono et al., 2023).
For GIST1M, with 09M and 10, the reported search curves of RNN-Descent are virtually identical to NSG and HNSW for recall up to 11. Construction times are given as approximately 12 s for NSG, approximately 13 s for NN-Descent, and approximately 14 s for RNN-Descent. The graph’s average out-degree is reported as approximately 15, stated to be the same order as NSG’s approximately 16–17, ensuring low memory cost (Ono et al., 2023).
These results support the characterization of RNN-Descent as a direct-construction ANNS graph that requires no prior 18-NN graph, needs no costly internal ANNS calls, converges in 19 time with small 20 and modest 21, and matches or exceeds the search performance of state-of-the-art graph indexes while cutting index build time by up to half (Ono et al., 2023).
5. GPU parallelization: GRNND
The later paper "GRNND: A GPU-Parallel Relative NN-Descent Algorithm for Efficient Approximate Nearest Neighbor Graph Construction" presents the first GPU-parallel algorithm of RNN-Descent designed to fully exploit GPU architecture (Li et al., 3 Oct 2025). Its starting point is the observation that as data amount and dimensionality increase, the complexity of graph construction in RNN-Descent rises sharply, making this stage increasingly time-consuming and even prohibitive for subsequent query processing.
The GRNND design introduces four mechanisms. First, disordered neighbor propagation replaces strict ascending-distance processing with random neighbor-pair sampling from 22, followed by RNG tests in arbitrary order; this is stated to mitigate synchronized update traps, enhance structural diversity, and avoid premature convergence during parallel execution (Li et al., 3 Oct 2025). Second, warp-level cooperative execution assigns one CUDA warp of 23 threads to each vertex update; distance computation uses striped loading over dimensions with warp-wide reduction via __shfl_down, while duplicate detection and insertion use __ballot and warp-max reductions rather than atomics or global locks (Li et al., 3 Oct 2025). Third, fixed-capacity double-buffered pools provide each vertex with two static arrays of length 24, one read set and one write set, which are swapped after each inner iteration; this removes dynamic allocation and lock contention (Li et al., 3 Oct 2025). Fourth, reverse-edge sampling adds only a fraction 25 of reverse edges, rather than all of them, to control memory growth and contention while maintaining structural diversity (Li et al., 3 Oct 2025).
In the GRNND formulation, one update has complexity 26, so over 27 vertices and 28 total refinement passes the total time is
29
with space complexity 30 (Li et al., 3 Oct 2025). The same source notes empirically that small 31 can suffice for high recall, citing the example 32 (Li et al., 3 Oct 2025).
Experiments compare GRNND with GPU baselines GANNS, CAGRA, and GGNN, and CPU baselines RNN-Descent(CPU) and HNSW on SIFT1M, DEEP1M, and GIST1M, using RTX 4090 and RTX 6000 Ada GPUs (Li et al., 3 Oct 2025). At Recall@10 near 33 on RTX 4090, construction times are reported as follows: on SIFT1M, GANNS 34 s, CAGRA 35 s, GGNN 36 s, RNN-Descent (CPU) 37 s, HNSW 38 s, and GRNND 39 s; on DEEP1M, the corresponding times are 40 s, 41 s, 42 s, 43 s, 44 s, and 45 s; on GIST1M, they are 46 s, 47 s, 48 s, 49 s, 50 s, and 51 s (Li et al., 3 Oct 2025). The reported speedups are 52–53 over existing GPU methods and 54–55 over CPU methods (Li et al., 3 Oct 2025).
Using a common CPU search routine over GPU-built indices, GRNND’s graphs are reported to achieve nearly identical Recall–QPS trade-offs to CPU RNN-Descent, while outperforming HNSW, GGNN, and CAGRA (Li et al., 3 Oct 2025). The ablations further state that ascending update order suffers convergence traps in parallel, descending order improves recall but at approximately 56 time cost, and disordered propagation gives the best balance, reaching recall near 57 with only approximately 58 the time of ascending order. For reverse-edge sampling, recall-versus-time curves on SIFT1M show an optimum at 59; low 60 builds fastest but yields recall around 61, while high 62 marginally raises recall to around 63 but costs approximately 64 more time. For iteration counts, the reported recommendation is 65 on SIFT1M, whereas on GIST1M values 66–67, 68 improve recall at moderate cost (Li et al., 3 Oct 2025).
6. Position within graph-based ANNS, distinctions, and open questions
RNN-Descent should be distinguished from two-stage refinement pipelines. The source explicitly describes it as a direct-construction algorithm that combines NN-Descent and RNG strategy so that useful edges are added and redundant edges pruned in the same iterative process, rather than by first materializing an approximate 69-NN graph and then refining it (Ono et al., 2023). That distinction is central to its reported construction-time advantage.
It should also be distinguished from the view that graph quality must be traded away to obtain faster construction. On GIST1M, the original paper reports search curves virtually identical to NSG and HNSW up to recall 70 while reducing construction time relative to both NSG and NN-Descent (Ono et al., 2023). The GPU study makes the same point in a different computational regime: GRNND is reported to deliver large construction-time speedups without loss of Search-Recall quality relative to CPU RNN-Descent (Li et al., 3 Oct 2025).
At the same time, the later work identifies explicit limitations and open directions. Fixed 71 may limit maximum recall in ultra-sparse regimes; multi-GPU or distributed extensions are described as non-trivial because of data-exchange patterns; and integration into billion-scale, out-of-core pipelines such as GPU clusters remains future work (Li et al., 3 Oct 2025). These statements delimit the current scope of the method family: the construction-time improvements are well established for the reported sparse-graph setting, but extension to broader hardware and larger-scale deployment remains an open research problem.
In this sense, RNN-Descent occupies a specific niche within graph-based ANNS: it is a sparse, direct-construction graph builder whose defining mechanism is the fusion of neighbor-of-neighbor proposal generation with RNG-based pruning, and whose later GPU realization preserves that structure while replacing sequential control flow and dynamic memory with randomized update order, warp-cooperative execution, double-buffered fixed-capacity pools, and sampled reverse-edge insertion (Ono et al., 2023).