Data-Free Dynamic Compression of CNNs for Tractable Efficiency
Abstract: To reduce the computational cost of convolutional neural networks (CNNs) on resource-constrained devices, structured pruning approaches have shown promise in lowering floating-point operations (FLOPs) without substantial drops in accuracy. However, most methods require fine-tuning or specific training procedures to achieve a reasonable trade-off between retained accuracy and reduction in FLOPs, adding computational overhead and requiring training data to be available. To this end, we propose HASTE (Hashing for Tractable Efficiency), a data-free, plug-and-play convolution module that instantly reduces a network's test-time inference cost without training or fine-tuning. Our approach utilizes locality-sensitive hashing (LSH) to detect redundancies in the channel dimension of latent feature maps, compressing similar channels to reduce input and filter depth simultaneously, resulting in cheaper convolutions. We demonstrate our approach on the popular vision benchmarks CIFAR-10 and ImageNet, where we achieve a 46.72% reduction in FLOPs with only a 1.25% loss in accuracy by swapping the convolution modules in a ResNet34 on CIFAR-10 for our HASTE module.
- Dimitris Achlioptas. Database-friendly random projections: Johnson-Lindenstrauss with binary coins. Journal of Computer and System Sciences, 66(4):671–687, 2003. ISSN 0022-0000. Special Issue on PODS 2001.
- Structured Pruning of Deep Convolutional Neural Networks. J. Emerg. Technol. Comput. Syst., 13(3), feb 2017. ISSN 1550-4832. doi: 10.1145/3005348.
- Batch-Shaping for Learning Conditional Channel Gated Networks. In International Conference on Learning Representations, 2020.
- SLIDE : In Defense of Smart Algorithms over Hardware Acceleration for Large-Scale Deep Learning Systems. In Proceedings of Machine Learning and Systems, 2020.
- MONGOOSE: A Learnable LSH Framework for Efficient Neural Network Training. In International Conference on Learning Representations, 2021.
- More is Less: A More Complicated Network with Less Inference Complexity. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1895–1903, Los Alamitos, CA, USA, jul 2017. IEEE Computer Society. doi: 10.1109/CVPR.2017.205.
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations, 2021.
- Fire Together Wire Together: A Dynamic Pruning Approach with Self-Supervised Mask Prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2022.
- Dynamic Channel Pruning: Feature Boosting and Suppression. In International Conference on Learning Representations, 2019.
- An Empirical Investigation of Catastrophic Forgeting in Gradient-Based Neural Networks. In Yoshua Bengio and Yann LeCun (eds.), 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Conference Track Proceedings, 2014.
- GhostNet: More Features From Cheap Operations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020.
- Deep Residual Learning for Image Recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770–778, 2016. doi: 10.1109/CVPR.2016.90.
- Filter Pruning via Geometric Median for Deep Convolutional Neural Networks Acceleration. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
- AMC: AutoML for Model Compression and Acceleration on Mobile Devices. In Vittorio Ferrari, Martial Hebert, Cristian Sminchisescu, and Yair Weiss (eds.), Computer Vision – ECCV 2018, pp. 815–832, Cham, 2018. Springer International Publishing. ISBN 978-3-030-01234-2.
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. 04 2017.
- Channel Gating Neural Networks. Curran Associates Inc., Red Hook, NY, USA, 2019.
- Approximate Nearest Neighbors: Towards Removing the Curse of Dimensionality. In Proceedings of the Thirtieth Annual ACM Symposium on Theory of Computing, STOC ’98, pp. 604–613, New York, NY, USA, 1998. Association for Computing Machinery. ISBN 0897919629. doi: 10.1145/276698.276876.
- Reformer: The Efficient Transformer. In International Conference on Learning Representations, 2020.
- Alex Krizhevsky. Learning Multiple Layers of Features from Tiny Images. 2009.
- Dynamic Dual Gating Neural Networks. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 5310–5319, 2021. doi: 10.1109/ICCV48922.2021.00528.
- Pruning Filters for Efficient ConvNets. In International Conference on Learning Representations, 2017.
- Very Sparse Random Projections. volume 2006, pp. 287–296, 08 2006. doi: 10.1145/1150402.1150436.
- Runtime Neural Pruning. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (eds.), Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017a.
- Towards Accurate Binary Convolutional Neural Network. In NIPS, 2017b.
- Dynamic Sparse Graph for Efficient Deep Learning. In International Conference on Learning Representations, 2019.
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 9992–10002, Los Alamitos, CA, USA, oct 2021. IEEE Computer Society. doi: 10.1109/ICCV48922.2021.00986.
- Learning Efficient Convolutional Networks through Network Slimming. In 2017 IEEE International Conference on Computer Vision (ICCV), pp. 2755–2763, Los Alamitos, CA, USA, oct 2017. IEEE Computer Society. doi: 10.1109/ICCV.2017.298.
- A ConvNet for the 2020s. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pp. 11966–11976. IEEE Computer Society, 2022. doi: 10.1109/CVPR52688.2022.01167.
- ThiNet: A Filter Level Pruning Method for Deep Neural Network Compression. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017, pp. 5068–5076. IEEE Computer Society, 2017. doi: 10.1109/ICCV.2017.541.
- ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design. In Proceedings of the European Conference on Computer Vision (ECCV), September 2018.
- Instant Neural Graphics Primitives with a Multiresolution Hash Encoding. ACM Trans. Graph., 41(4):102:1–102:15, July 2022. doi: 10.1145/3528223.3530127.
- PyTorch: An Imperative Style, High-Performance Deep Learning Library, 2019.
- Huy Phan. huyvnphan/pytorch_cifar10, January 2021. URL https://doi.org/10.5281/zenodo.4431043.
- Memory-Efficient Implementation of DenseNets. CoRR, abs/1707.06990, 2017.
- ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision, 115(3):211–252, December 2015. ISSN 0920-5691. doi: 10.1007/s11263-015-0816-y.
- MobileNetV2: Inverted Residuals and Linear Bottlenecks. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4510–4520, Los Alamitos, CA, USA, jun 2018. IEEE Computer Society. doi: 10.1109/CVPR.2018.00474.
- Very Deep Convolutional Networks for Large-Scale Image Recognition. In International Conference on Learning Representations, 2015.
- EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In Kamalika Chaudhuri and Ruslan Salakhutdinov (eds.), Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pp. 6105–6114. PMLR, 09–15 Jun 2019.
- EfficientNetV2: Smaller Models and Faster Training. In Marina Meila and Tong Zhang (eds.), Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp. 10096–10106. PMLR, 18–24 Jul 2021.
- Dynamic Convolutions: Exploiting Spatial Sparsity for Faster Inference. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2317–2326, Los Alamitos, CA, USA, jun 2020. IEEE Computer Society. doi: 10.1109/CVPR42600.2020.00239.
- Learning Structured Sparsity in Deep Neural Networks. In Proceedings of the 30th International Conference on Neural Information Processing Systems, NIPS’16, pp. 2082–2090. Curran Associates Inc., 2016. ISBN 9781510838819.
- Dimensionality reduced training by pruning and freezing parts of a deep neural network: a survey. Artificial Intelligence Review, 2023. ISSN 1573-7462. doi: 10.1007/s10462-023-10489-1.
- ConvNeXt V2: Co-Designing and Scaling ConvNets With Masked Autoencoders. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 16133–16142, June 2023.
- An Efficient Channel-level Pruning for CNNs without Fine-tuning. In 2021 International Joint Conference on Neural Networks (IJCNN), pp. 1–8, 2021. doi: 10.1109/IJCNN52387.2021.9533397.
- LookupFFN: Making Transformers Compute-lite for CPU inference. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (eds.), Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pp. 40707–40718. PMLR, 23–29 Jul 2023.
- Accelerating Very Deep Convolutional Networks for Classification and Detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 38(10):1943–1955, oct 2016. ISSN 1939-3539. doi: 10.1109/TPAMI.2015.2502579.
- Shufflenet: An extremely efficient convolutional neural network for mobile devices. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018.
- Discrimination-Aware Channel Pruning for Deep Neural Networks. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, NIPS’18, pp. 883–894, Red Hook, NY, USA, 2018. Curran Associates Inc.
Paper Prompts
Sign up for free to create and run prompts on this paper.