How much pre-training is enough to discover a good subnetwork? (2108.00259v3)

Published 31 Jul 2021 in stat.ML, cs.AI, cs.LG, and math.OC

Abstract: Neural network pruning is useful for discovering efficient, high-performing subnetworks within pre-trained, dense network architectures. More often than not, it involves a three-step process -- pre-training, pruning, and re-training -- that is computationally expensive, as the dense model must be fully pre-trained. While previous work has revealed through experiments the relationship between the amount of pre-training and the performance of the pruned network, a theoretical characterization of such dependency is still missing. Aiming to mathematically analyze the amount of dense network pre-training needed for a pruned network to perform well, we discover a simple theoretical bound in the number of gradient descent pre-training iterations on a two-layer, fully-connected network, beyond which pruning via greedy forward selection [61] yields a subnetwork that achieves good training error. Interestingly, this threshold is shown to be logarithmically dependent upon the size of the dataset, meaning that experiments with larger datasets require more pre-training for subnetworks obtained via pruning to perform well. Lastly, we empirically validate our theoretical results on a multi-layer perceptron trained on MNIST.

Citations (3)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Related Papers

Pruning Before Training May Improve Generalization, Provably (2023)
When to Prune? A Policy towards Early Structural Pruning (2021)
Good Subnetworks Provably Exist: Pruning via Greedy Forward Selection (2020)
Robust Pruning at Initialization (2020)
Pruning from Scratch (2019)