Papers
Topics
Authors
Recent
Search
2000 character limit reached

N-ImageNetV2: Multi-Label Benchmarking

Updated 13 July 2026
  • N-ImageNetV2 is a multi-label evaluation protocol that assigns variable top-k predictions corresponding to the number of valid labels per image.
  • The approach reveals that the classic top-1 accuracy gap between ImageNet and ImageNetV2 largely stems from the restrictive single-label evaluation.
  • By incorporating metrics like ReaL accuracy and ASMA, N-ImageNetV2 offers a more balanced assessment of model performance under label multiplicity.

N-ImageNetV2 denotes, in the formulation developed in "The Impact of the Single-Label Assumption in Image Recognition Benchmarking," an ImageNetV2 evaluation setting in which each image is associated with NN valid labels and model predictions are evaluated with variable top-kk where k=Nk=N, rather than under a single-label top-1 protocol (Anzaku et al., 2024). The formulation is motivated by the mismatch between the single-label assumption used in standard ImageNet-style benchmarking and the multi-label reality of many natural images. In that view, the widely discussed ImageNet-to-ImageNetV2 accuracy gap is substantially affected by evaluation misalignment. A distinct but adjacent line of work on dataset cleanup does not define a variant called “N-ImageNetV2,” but does release a cleaned ImageNetV2 MatchedFrequency split, denoted V2cV2_c, which is relevant when “N” is informally taken to mean a noise-reduced ImageNetV2 (Kertész, 2021).

1. Definition and conceptual scope

In the multi-label formulation, the core idea behind N-ImageNetV2 is to evaluate models on ImageNetV2 with NN valid labels per image and to set k=Nk=N when constructing predictions from a model’s ranked outputs (Anzaku et al., 2024). The resulting protocol is not merely a change from top-1 to top-kk in the conventional sense; it treats kk as image-dependent and determined by the number of valid labels attached to each image.

This definition is tightly connected to the paper’s broader notion of Multi-Label Prediction Capability (MLPC), which asks how well single-label-trained ImageNet models rank multiple correct categories, not just whether the top prediction matches a single annotated class (Anzaku et al., 2024). In this framing, N-ImageNetV2 is best understood as an evaluation protocol for ImageNetV2 that is explicitly sensitive to label multiplicity.

A separate interpretation arises from automated dataset-cleaning work. "Automated Cleanup of the ImageNet Dataset by Model Consensus, Explainability and Confident Learning" states that it does not define a variant called “N-ImageNetV2”; the closest artifact is its cleaned ImageNetV2 MatchedFrequency validation split, V2cV2_c (Kertész, 2021). This suggests that the term is presently better grounded as a multi-label evaluation concept than as a canonical cleaned dataset release.

2. Label multiplicity in ImageNetV2

The rationale for N-ImageNetV2 depends on an observed asymmetry between ImageNetV1 and ImageNetV2 in label multiplicity. Using refined annotations, the multi-label study reports that roughly 16% of ImageNetV1 images have more than one valid label, versus about 48% in ImageNetV2 (Anzaku et al., 2024). The paper gives the corresponding counts from Table 1: for ImageNetV1 with ReaL labels over 46,837 annotated images, 39,394 have exactly one label and 7,443 have at least two labels; for ImageNetV2 over 9,858 images, 5,083 have exactly one label and 4,775 have at least two labels.

These differences matter because standard top-1 evaluation compares a model’s first-ranked prediction against a single designated class. When an image admits multiple valid categories, a model can produce a semantically correct answer that is treated as an error if that label is not the one annotated under the single-label protocol. The multi-label study argues that ImageNetV2’s higher multiplicity increases the frequency of such valid-but-unannotated predictions, thereby inflating the apparent degradation under standard scoring (Anzaku et al., 2024).

This interpretation directly reframes the historical ImageNetV2 result. Prior work reported an 11–14% top-1 accuracy drop on ImageNetV2 relative to ImageNet. The multi-label analysis argues that a major and previously understudied factor behind this decline is the higher proportion of multi-label images in ImageNetV2, and that once multi-label characteristics are taken into account, the apparent degradation largely disappears (Anzaku et al., 2024). A plausible implication is that N-ImageNetV2 is intended to separate genuine distribution-shift effects from artifacts introduced by restrictive evaluation.

3. Variable top-kk and MLPC-aware metrics

The operational definition of N-ImageNetV2 begins with a variable top-kk0 prediction rule. Let the label space be kk1. For image kk2, let the set of valid labels be kk3 and let kk4. If a model produces ranked predictions kk5, the predicted set is defined as

kk6

The paper lists two image-level correctness criteria that can then be applied:

kk7

and

kk8

The corresponding dataset-level accuracies are

kk9

In the paper’s primary implementation, however, variable top-k=Nk=N0 is used to build a binary multi-label prediction vector k=Nk=N1 by setting k=Nk=N2 if k=Nk=N3 and k=Nk=N4 otherwise, with k=Nk=N5 for ImageNet-1K (Anzaku et al., 2024).

MLPC is then assessed with three metrics. The first is standard top-1 accuracy,

k=Nk=N6

The second is ReaL accuracy, which accepts multiple plausible labels per image but still checks only the top-1 prediction:

k=Nk=N7

The third is Aggregate Subgroup Model Accuracy (ASMA), which is computed from the variable top-k=Nk=N8 multi-label predictions and aggregates performance across subgroups defined by label count (Anzaku et al., 2024).

ASMA is designed to control for imbalance across images with different numbers of valid labels. If subgroups k=Nk=N9 partition the dataset by label multiplicity, with V2cV2_c0 images in subgroup V2cV2_c1, the per-example label-wise accuracy is

V2cV2_c2

the subgroup accuracy is

V2cV2_c3

and ASMA with uniform subgroup weighting is

V2cV2_c4

The paper reports ASMA on datapoints with label counts from 1 to 5 in order to ensure comparability across ImageNetV1 and ImageNetV2 (Anzaku et al., 2024).

The significance of this metric suite is methodological. ReaL relaxes the single-label assumption only partially because it still evaluates only the first-ranked prediction. ASMA and variable top-V2cV2_c5 are intended to recover information about how many valid labels a model ranks highly, which conventional top-1 evaluation discards (Anzaku et al., 2024).

4. Empirical behavior across 315 ImageNet-trained models

The multi-label study evaluates 315 TIMM checkpoints spanning CNNs and Transformers. Under conventional top-1 accuracy, the models exhibit a visible drop from ImageNetV1 to ImageNetV2, consistent with prior reports. Under ReaL accuracy and ASMA, however, the gap narrows substantially and is often near zero or positive (Anzaku et al., 2024).

Model Top-1 V1 V2cV2_c6 V2 ReaL / ASMA V1 V2cV2_c7 V2
EVA-L/14 336 (in22k ft in1k) 88.50% V2cV2_c8 80.99% ReaL 90.76% V2cV2_c9 91.74%; ASMA 72.22 NN0 70.35
VOLO-D5 512 (in1k) 87.04% NN1 77.87% ReaL 90.53% NN2 92.03%; ASMA 72.34 NN3 73.40
ConvNeXtV2-Huge (in22k→in1k 512) 88.84% NN4 80.52% ReaL 90.56% NN5 91.05%; ASMA 68.52 NN6 68.98

These examples illustrate the paper’s central claim. Top-1 gaps are large, often on the order of 7–10 percentage points for high-performing models in the reported examples, but ReaL and ASMA gaps are near-zero and frequently positive (Anzaku et al., 2024). The reported pattern is that ImageNetV2 often appears worse only when evaluation is constrained to a single designated label.

The aggregate analyses strengthen that interpretation. Figures 2 and 3 reportedly show that top-1 differences are consistently negative for ImageNetV2, whereas both ReaL accuracy and ASMA substantially narrow the gap and often eliminate it, especially for images with at least two labels (Anzaku et al., 2024). The subgroup boxplots further indicate that median MLPC on ImageNetV2 is comparable to or higher than that on ImageNetV1 for label counts of at least three.

The paper also emphasizes that MLPC reveals model differences not captured by top-1 accuracy. VOLO models are described as consistently placing multiple valid labels high in the ranking, leading to stronger ASMA and better variable top-NN7 behavior (Anzaku et al., 2024). This is important for N-ImageNetV2 because the protocol is explicitly intended to rank models by their capacity to recover multiple correct labels, not only by their first guess.

5. PatchML and controlled stress testing

To further isolate multi-label recognition from contextual cues, the paper introduces PatchML, a synthetic multi-label test set built by extracting object patches from ImageNet Object Localization annotations and compositing NN8 patches per image for NN9 onto a k=Nk=N0 blank canvas without overlap (Anzaku et al., 2024). The generation procedure has two stages: patch extraction and aggregation to create a labeled patch pool k=Nk=N1, followed by multi-label image generation in which k=Nk=N2 patches are sampled without replacement, resized to a fixed patch size k=Nk=N3 that depends on k=Nk=N4, placed on a grid with random offsets, and assigned the union of their labels as ground truth:

k=Nk=N5

Across five seeds, the dataset sizes per seed are reported as roughly 26–27k images with 2 labels, 17–18k with 3 labels, 13k with 4, 8.7k with 6, and 5.7k with 9 labels (Anzaku et al., 2024). PatchML is therefore not a replacement for ImageNetV2, but a stress test for the same underlying question: can models trained with single-label supervision nonetheless recognize multiple objects?

The reported answer is affirmative, though highly model-dependent. Across the same 315 models, subgroup accuracy curves on PatchML show broad variability and identify architectures that are comparatively strong at recognizing multiple objects even when trained on single-label data (Anzaku et al., 2024). The top models by PatchML ASMA include EVA-L/14 336 in22k→in1k with 74.50, ConvNeXtV2-Huge fcmae in22k→in1k 512 with 71.45, VOLO-D5 448/512 with approximately 70.2–70.3, and CAiT-M48 448 with 67.11.

PatchML also induces substantial reordering relative to ImageNetV1 top-1 ranking. CAiT-M48 448 improves by 71 positions, VOLO-D5 448/512 by 53, and EVA-L/14 336 by 13 when ranked by PatchML ASMA rather than ImageNetV1 top-1 (Anzaku et al., 2024). This supports the paper’s conclusion that single-label top-1 evaluation undervalues the capability to recognize multiple objects and motivates N-ImageNetV2-style reporting beyond standard classification accuracy.

6. Relation to cleaned ImageNetV2 and naming ambiguity

The term “N-ImageNetV2” is not used uniformly across the relevant literature. The automated cleanup paper explicitly states that it does not define a variant called “N-ImageNetV2” and that its contribution on the ImageNetV2 side is a cleaned MatchedFrequency validation split, k=Nk=N6 (Kertész, 2021). That split is produced through an automated pipeline combining confident learning, top-5 model consensus, and, where bounding boxes exist, CAM-based exemption. Because ImageNetV2 lacks bounding boxes in that setup, the explainability exemption is skipped for k=Nk=N7 (Kertész, 2021).

For ImageNetV2 MatchedFrequency, the reported cleaning outcomes are 292 label fixes and 524 removals, corresponding to 2.9% and 5.2% respectively (Kertész, 2021). The paper evaluates only the MatchedFrequency split; TopImages and Threshold0.7 are not covered. It reports consistent performance gains on k=Nk=N8 relative to the original k=Nk=N9 for several models. For example, EfficientNet-B0 evaluated on kk0 yields 61.53/82.23 top-1/top-5, whereas on kk1 it yields 68.58/88.87; SqueezeNet 1.1 improves from 42.69/65.15 to 47.72/71.38 (Kertész, 2021).

This cleaned-split work addresses label noise and ambiguity, whereas the N-ImageNetV2 proposal in the multi-label paper addresses evaluation misalignment induced by label multiplicity. The two are complementary rather than interchangeable. A plausible implication is that a future benchmark could combine both ideas: use a cleaned ImageNetV2 split while also evaluating with variable top-kk2 and MLPC-aware metrics.

Several limitations constrain current N-ImageNetV2-style evaluation. First, the multi-label annotations themselves are not exhaustive. The paper notes that ReaL labels for ImageNetV1 and Multilabelfy annotations for ImageNetV2 are subject to ambiguity from synonymy, fine-grained distinctions, and taxonomy overlaps (Anzaku et al., 2024). Second, ASMA is restricted in the reported experiments to images with 1–5 labels for cross-dataset comparability, leaving behavior at very high label counts underexplored.

Third, ASMA uses per-class equality across kk3 classes, so true negatives dominate and may inflate absolute values, even though subgrouping mitigates imbalance (Anzaku et al., 2024). The paper therefore mentions example-based alternatives such as per-image Jaccard accuracy over the union of predicted and valid labels. It also identifies rank-sensitive MLPC metrics—such as the expected rank of the second or third valid label—as an open direction not covered by the study.

Fourth, all 315 evaluated models were trained under single-label supervision. The paper explicitly notes that extending the analysis to Single Positive Multi-label Learning (SPML) or to multi-label training could improve MLPC and alter model rankings (Anzaku et al., 2024). Thus, the current results characterize the latent multi-label competence of standard ImageNet models rather than an upper bound on multi-label recognition.

The practical reporting recommendations are correspondingly explicit. For N-ImageNetV2, the paper recommends variable top-kk4 prediction with kk5, MLPC metrics beyond top-1 including ReaL accuracy and ASMA, subgroup reporting by label multiplicity, and publication of subgroup counts kk6 and the distribution of kk7 (Anzaku et al., 2024). It also recommends continuing to report top-1 and top-5 for continuity, but interpreting them cautiously on multi-label images. In that formulation, the central conclusion is that when ImageNetV2 is evaluated with the number of valid labels per image and with multi-label-aware metrics, the ImageNet-to-ImageNetV2 “accuracy gap” largely reflects artifacts of single-label scoring rather than a pure failure of generalization (Anzaku et al., 2024).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to N-ImageNetV2.