Block-Sparse Recurrent Neural Networks (1711.02782v1)

Published 8 Nov 2017 in cs.LG, cs.AI, and stat.ML

Abstract: Recurrent Neural Networks (RNNs) are used in state-of-the-art models in domains such as speech recognition, machine translation, and LLMling. Sparsity is a technique to reduce compute and memory requirements of deep learning models. Sparse RNNs are easier to deploy on devices and high-end server processors. Even though sparse operations need less compute and memory relative to their dense counterparts, the speed-up observed by using sparse operations is less than expected on different hardware platforms. In order to address this issue, we investigate two different approaches to induce block sparsity in RNNs: pruning blocks of weights in a layer and using group lasso regularization to create blocks of weights with zeros. Using these techniques, we demonstrate that we can create block-sparse RNNs with sparsity ranging from 80% to 90% with small loss in accuracy. This allows us to reduce the model size by roughly 10x. Additionally, we can prune a larger dense network to recover this loss in accuracy while maintaining high block sparsity and reducing the overall parameter count. Our technique works with a variety of block sizes up to 32x32. Block-sparse RNNs eliminate overheads related to data storage and irregular memory accesses while increasing hardware efficiency compared to unstructured sparsity.

PDF Abstract

Summarize PDF Markdown Bookmark Chat (Pro)

Authors (3)

Sharan Narang (31 papers)
Eric Undersander (11 papers)
Gregory Diamos (11 papers)

Citations (133)

View on Semantic Scholar

Block-Sparse Recurrent Neural Networks (1711.02782v1)

Related Papers