Papers
Topics
Authors
Recent
Search
2000 character limit reached

An Analysis of Self-supervised Pre-training with Dependent Samples

Published 4 Sep 2026 in stat.ML and cs.LG | (2609.05031v1)

Abstract: Self-supervised learning relies on so-called data augmentations φ(x)φ(x) of unlabeled datapoints xx --- for example, masking random pixels in an image xx --- that should leave the label of xx invariant and are often used to learn a lower-complexity invariant subspace V\cal V for downstream tasks. In practice, such augmentations φl(xi){ φ_l(x_i) } are pooled together to learn V\cal V, despite obvious inter-dependencies between different augmentations φl(x),φk(x)φ_l(x), φ_k(x) of the same datapoint xx. However, theoretical works on the subject typically consider procedures that avoid such dependencies, and are therefore limited to operate on smaller subsets of independent data. We show in this work that pooling augmentations together, despite inter-dependencies, is a better alternative than the baseline of partitioning the data into subsets of independent data. More precisely, in the context of estimating V\cal V, the statistical estimation error bounds for pooling are never worse than the partitioning baseline, and in some cases --- such as masking or noise injection-based augmentations over a shallow neural network --- naive pooling leads to faster rates in terms of the number of augmentations. The benefits of pooling are particularly prominent when the correlations between different augmentations φl(x),φk(x)φ_l(x), φ_k(x) have mild effects on estimation or help decrease the estimation variance. The analysis, therefore, yields new insights into the success of pooling augmented samples in self-supervised pre-training, and provides an intuition behind the practical preference towards using many augmentations.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 2 likes about this paper.