2000 character limit reached
MLonMCU: TinyML Benchmarking with Fast Retargeting
Published 15 Jun 2023 in cs.LG | (2306.08951v1)
Abstract: While there exist many ways to deploy machine learning models on microcontrollers, it is non-trivial to choose the optimal combination of frameworks and targets for a given application. Thus, automating the end-to-end benchmarking flow is of high relevance nowadays. A tool called MLonMCU is proposed in this paper and demonstrated by benchmarking the state-of-the-art TinyML frameworks TFLite for Microcontrollers and TVM effortlessly with a large number of configurations in a low amount of time.
- R. David, J. Duke, A. Jain, V. Janapa Reddi, N. Jeffries, J. Li, N. Kreeger, I. Nappier, M. Natraj, T. Wang et al., “Tensorflow lite micro: Embedded machine learning for tinyml systems,” Proceedings of Machine Learning and Systems, vol. 3, pp. 800–811, 2021.
- M. S. Louis, Z. Azad, L. Delshadtehrani, S. Gupta, P. Warden, V. J. Reddi, and A. Joshi, “Towards deep learning using tensorflow lite on risc-v,” in Third Workshop on Computer Architecture Research with RISC-V (CARRV), vol. 1, 2019, p. 6.
- R. Stahl, ““exploring static code generation and simd-acceleration for machine learning on risc-v” in in “risc-v forum: Developer tools & tool chains”,” 2021, [Recording available: https://youtu.be/NLGAjdVIzkk]. [Online]. Available: https://riscvforumdttc2021.sched.com/event/jGkT
- T. Chen, T. Moreau, Z. Jiang, L. Zheng, E. Yan, H. Shen, M. Cowan, L. Wang, Y. Hu, L. Ceze et al., “{{\{{TVM}}\}}: An automated {{\{{End-to-End}}\}} optimizing compiler for deep learning,” in 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18), 2018, pp. 578–594.
- C. Banbury, V. J. Reddi, P. Torelli, J. Holleman, N. Jeffries, C. Kiraly, P. Montino, D. Kanter, S. Ahmed, D. Pau et al., “Mlperf tiny benchmark,” arXiv preprint arXiv:2106.07597, 2021.
- D. Mueller-Gritschneder, M. Dittrich, M. Greim, K. Devarajegowda, W. Ecker, and U. Schlichtmann, “The extendable translating instruction set simulator (etiss) interlinked with an mda framework for fast risc prototyping,” in 2017 International Symposium on Rapid System Prototyping (RSP). IEEE, 2017, pp. 79–84.
- L. Lai, N. Suda, and V. Chandra, “Cmsis-nn: Efficient neural network kernels for arm cortex-m cpus,” arXiv preprint arXiv:1801.06601, 2018.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.
Top Community Prompts
Collections
Sign up for free to add this paper to one or more collections.