26ms Inference Time for ResNet-50: Towards Real-Time Execution of all DNNs on Smartphone (1905.00571v1)

Published 2 May 2019 in cs.LG, cs.CV, and stat.ML

Abstract: With the rapid emergence of a spectrum of high-end mobile devices, many applications that required desktop-level computation capability formerly can now run on these devices without any problem. However, without a careful optimization, executing Deep Neural Networks (a key building block of the real-time video stream processing that is the foundation of many popular applications) is still challenging, specifically, if an extremely low latency or high accuracy inference is needed. This work presents CADNN, a programming framework to efficiently execute DNN on mobile devices with the help of advanced model compression (sparsity) and a set of thorough architecture-aware optimization. The evaluation result demonstrates that CADNN outperforms all the state-of-the-art dense DNN execution frameworks like TensorFlow Lite and TVM.

PDF Abstract

Summarize PDF Markdown Bookmark Chat (Pro)

Authors (4)

Wei Niu (68 papers)
Xiaolong Ma (57 papers)
Yanzhi Wang (197 papers)
Bin Ren (136 papers)

Citations (22)

View on Semantic Scholar

26ms Inference Time for ResNet-50: Towards Real-Time Execution of all DNNs on Smartphone (1905.00571v1)

Related Papers