FLAASH: Flexible Accelerator Architecture for Sparse High-Order Tensor Contraction (2404.16317v1)

Published 25 Apr 2024 in cs.AR and cs.LG

Abstract: Tensors play a vital role in ML and often exhibit properties best explored while maintaining high-order. Efficiently performing ML computations requires taking advantage of sparsity, but generalized hardware support is challenging. This paper introduces FLAASH, a flexible and modular accelerator design for sparse tensor contraction that achieves over 25x speedup for a deep learning workload. Our architecture performs sparse high-order tensor contraction by distributing sparse dot products, or portions thereof, to numerous Sparse Dot Product Engines (SDPEs). Memory structure and job distribution can be customized, and we demonstrate a simple approach as a proof of concept. We address the challenges associated with control flow to navigate data structures, high-order representation, and high-sparsity handling. The effectiveness of our approach is demonstrated through various evaluations, showcasing significant speedup as sparsity and order increase.

References (21)

Citations (1)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Tweets

https://twitter.com/WWVY/status/1783711602878902660

FLAASH: Flexible Accelerator Architecture for Sparse High-Order Tensor Contraction (2404.16317v1)

Summary

Related Papers

Tweets