Papers
Topics
Authors
Recent
Search
2000 character limit reached

MAP@100: Ranking Quality Metric

Updated 9 March 2026
  • MAP@100 is a metric that quantifies ranking quality by calculating average precision over only the top 100 results, integrating both relevance and rank order.
  • It is applied in scenarios such as query evaluation, large-scale image retrieval, and deep learning pipelines by leveraging differentiable surrogates like smooth histogram-binning.
  • Analytic baselines and statistical significance tests using MAP@100 provide a rigorous framework for benchmarking and validating improvements in ranking algorithms.

Mean Average Precision at 100 (MAP@100) is a widely used metric in information retrieval and recommender systems, designed to evaluate the quality of ranking algorithms when only the top 100 results are of interest. MAP@100 integrates both the relevance and ranking positions of retrieved items, and its application spans query-based evaluation, large-scale image retrieval, and statistical significance testing of ranking improvements (Manzhos et al., 4 Nov 2025, Revaud et al., 2019).

1. Formal Definitions and Notational Framework

Let NN denote the total number of candidate items and RR the subset of relevant items (R≤NR\leq N). For a retrieved list truncated at position k=100k=100:

  • rel(i)∈{0,1}\mathrm{rel}(i)\in\{0,1\} indicates relevance of item at position ii.
  • P@i=1i∑j=1irel(j)\mathrm{P}@i = \frac{1}{i} \sum_{j=1}^i \mathrm{rel}(j) is the precision at position ii.
  • M=min⁡(R,100)M=\min(R,100) is the normalization denominator.
  • The kkth harmonic number is RR0; RR1.

For a single ranking, the Average Precision at 100 is

RR2

For a set RR3 of RR4 queries or users, each with its own RR5,

RR6

This definition ensures that MAP@100 is sensitive to both the number and distribution of relevant items within the top 100 ranks (Manzhos et al., 4 Nov 2025, Revaud et al., 2019).

2. Random Baseline: Expectation and Variance under Uniform Rankings

MAP@100’s significance is enhanced by analytic baselines under the random-ranking model, where RR7 relevant items are uniformly distributed among RR8 candidates (sampling without replacement).

The expectation of RR9 is

R≤NR\leq N0

where R≤NR\leq N1 (Manzhos et al., 4 Nov 2025). This establishes the expected MAP@100 achievable by chance.

The variance has the form

R≤NR\leq N2

with explicit expressions for R≤NR\leq N3 as functions of R≤NR\leq N4 and R≤NR\leq N5. For R≤NR\leq N6 independent queries,

R≤NR\leq N7

and for homogeneous queries R≤NR\leq N8,

R≤NR\leq N9

These baselines are fundamental for statistical testing and for contextualizing observed system performance (Manzhos et al., 4 Nov 2025).

3. MAP@100 for Batched Image Retrieval and Listwise Optimization

In deep image retrieval systems, MAP@100 is computed as follows. Let k=100k=1000 be normalized descriptors and k=100k=1001 the cosine similarity. The database k=100k=1002 is searched for relevant items corresponding to query k=100k=1003. The definitions proceed as:

  • k=100k=1004: relevance label for query k=100k=1005, database item k=100k=1006.
  • k=100k=1007: total relevant images.

Truncated average precision at k=100k=1008,

k=100k=1009

with rel(i)∈{0,1}\mathrm{rel}(i)\in\{0,1\}0 and rel(i)∈{0,1}\mathrm{rel}(i)\in\{0,1\}1 as the precision and recall increments at position rel(i)∈{0,1}\mathrm{rel}(i)\in\{0,1\}2. Over rel(i)∈{0,1}\mathrm{rel}(i)\in\{0,1\}3 queries,

rel(i)∈{0,1}\mathrm{rel}(i)\in\{0,1\}4

This formulation aligns with practical retrieval and learning pipelines (Revaud et al., 2019).

4. Differentiable Surrogates: Smooth Histogram-Binning for AP@100

Classic AP@100 calculation is non-differentiable due to explicit sorting and indicator functions. A smooth surrogate is constructed by soft-binning the score axis rel(i)∈{0,1}\mathrm{rel}(i)\in\{0,1\}5 into rel(i)∈{0,1}\mathrm{rel}(i)\in\{0,1\}6 bins of width rel(i)∈{0,1}\mathrm{rel}(i)\in\{0,1\}7, each centered at rel(i)∈{0,1}\mathrm{rel}(i)\in\{0,1\}8. The kernel

rel(i)∈{0,1}\mathrm{rel}(i)\in\{0,1\}9

provides a soft assignment, and soft histograms for all and relevant items are accumulated:

  • ii0,
  • ii1.

Approximated precision and recall per bin are

ii2

yielding quantized AP,

ii3

This AP surrogate is differentiable w.r.t. network parameters, enabling direct end-to-end optimization (Revaud et al., 2019).

5. Computation, Training, and Memory Considerations

Sorting-based ii4 requires ii5 per query. The histogram-binning approximation bypasses explicit sorting with computational cost ii6 per query. For batched training, ii7 (batch size), yielding ii8 operations per batch. Memory usage is dominated by the ii9 similarity matrix and P@i=1i∑j=1irel(j)\mathrm{P}@i = \frac{1}{i} \sum_{j=1}^i \mathrm{rel}(j)0 descriptors. For example, with P@i=1i∑j=1irel(j)\mathrm{P}@i = \frac{1}{i} \sum_{j=1}^i \mathrm{rel}(j)1 and descriptor dimension P@i=1i∑j=1irel(j)\mathrm{P}@i = \frac{1}{i} \sum_{j=1}^i \mathrm{rel}(j)2, total memory footprint is P@i=1i∑j=1irel(j)\mathrm{P}@i = \frac{1}{i} \sum_{j=1}^i \mathrm{rel}(j)3—well within typical GPU memory budgets (Revaud et al., 2019).

Training with this surrogate involves:

  • Forward-passing all P@i=1i∑j=1irel(j)\mathrm{P}@i = \frac{1}{i} \sum_{j=1}^i \mathrm{rel}(j)4 images to obtain descriptors.
  • Computing the similarity matrix, AP surrogates, and loss gradients wrt descriptors (P@i=1i∑j=1irel(j)\mathrm{P}@i = \frac{1}{i} \sum_{j=1}^i \mathrm{rel}(j)5 memory/compute).
  • Backpropagating gradients by recomputing each image’s forward pass individually, eliminating the need to store all activations. This staged procedure optimally utilizes GPU resources and provides 2–3P@i=1i∑j=1irel(j)\mathrm{P}@i = \frac{1}{i} \sum_{j=1}^i \mathrm{rel}(j)6 speed-ups over alternative approaches such as hard-negative mining (Revaud et al., 2019).

6. Statistical Significance and Null Model Interpretations

MAP@100 values are conventionally interpreted relative to random-ranking baselines. Compute the mean (P@i=1i∑j=1irel(j)\mathrm{P}@i = \frac{1}{i} \sum_{j=1}^i \mathrm{rel}(j)7) and standard deviation (P@i=1i∑j=1irel(j)\mathrm{P}@i = \frac{1}{i} \sum_{j=1}^i \mathrm{rel}(j)8) for the random model. Given an observed P@i=1i∑j=1irel(j)\mathrm{P}@i = \frac{1}{i} \sum_{j=1}^i \mathrm{rel}(j)9, the standardized ii0-score is:

ii1

Under the null (random) hypothesis, ii2 is approximately standard normal. ii3 implies statistical significance at ii4. This framework enables researchers to rigorously assess if observed ranking gains exceed those explainable by chance, with analytic baselines for mean and fluctuation scale (Manzhos et al., 4 Nov 2025).

7. Practical Relevance and Context Among Metrics

MAP@100 is preferred in scenarios where only the top-ranked results are critical, such as web search, recommender system outputs, and image retrieval tasks. Compared to untruncated MAP, MAP@100 more closely models user-facing scenarios where lower-ranked results are rarely examined. The differentiable surrogates developed for deep learning pipelines facilitate direct optimization of retrieval objectives, outperforming proxy loss functions or heuristic approaches (Revaud et al., 2019). The closed-form random baselines further enhance MAP@100’s interpretability and robustness for benchmarking systems (Manzhos et al., 4 Nov 2025).


References:

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Mean Average Precision at 100 (MAP@100).