Determine the cause of DINOv2’s performance advantage

Determine whether DINOv2’s superior retrieval performance over supervised ImageNet representations is attributable solely to self-supervised pretraining or also to the difference in pretraining data.

Background

The paper reports that the self-supervised DINOv2 ViT-S/14 backbone substantially outperforms the supervised ImageNet-pretrained ViT-S/16 backbone across cropping conditions. However, the two models differ in both pretraining objective and pretraining dataset: DINOv2 uses LVD-142M, whereas the supervised model uses ImageNet-1k. Because these factors are confounded, the paper does not establish which factor causes the observed performance gain and identifies determining the actual cause as future work.

References

Additionally, we found that DINOv2 outperforms standard supervised representations, though we do not fully establish whether this is due to self-supervision alone or also to the difference in training data; in future work, we aim to pinpoint the actual cause.

— Automated Goldsmith's Mark Retrieval in Silverware  (2609.20509 - Tiwari et al., 17 Sep 2026) in Section 4, Conclusion