misi: a Metric Inverted Sample Index
Abstract: We present misi, an inverted index for approximate nearest-neighbor search over general metric spaces whose vocabulary is a random sample of the database, of size proportional to . Each object is represented by its nearest sample points, found by a pluggable inner index over the sample; queries are answered by an idf-weighted shared-neighbor vote followed by exact verification of candidates. The construction generalizes the NAPP index from a constant number of pivots to a linear-size vocabulary, which keeps posting lists at constant expected length as grows and turns the index into a combinator: any high-recall index on points yields an index on points, for any metric. A probabilistic model gives a recall guarantee -- logarithmic in over the overlap gap suffices, with a verification budget the index itself estimates -- and a matching limit: the vote cannot resolve overlap differences below order . The design's strengths are structural: construction is independent searches -- embarrassingly parallel, deterministic, s for $108$ vectors on 64 cores, faster than a matched-recall graph build -- it streams under an enforced 3 GiB cap, and the portable artifact serves $108$ vectors from NVMe within an enforced 8 GB budget, below the working floor of the SSD-graph baseline. Its cost is query-time work: saturated graph baselines answer $6$- faster in RAM, and the verification budget for 0.99 recall grows as . All results carry seeds, saturation sweeps and full configurations, are generated from run manifests, and include measured negative results. The intended applications weight construction cost, determinism, memory footprint, or black-box metrics over peak throughput: frequently rebuilt corpora, batch similarity workloads, constrained-memory serving.
Paper Prompts
Sign up for free to create and run prompts on this paper.