Efficiently Distilling LLMs for Edge Applications (2404.01353v1)

Published 1 Apr 2024 in cs.LG, cs.AI, and cs.CL

Abstract: Supernet training of LLMs is of great interest in industrial applications as it confers the ability to produce a palette of smaller models at constant cost, regardless of the number of models (of different size / latency) produced. We propose a new method called Multistage Low-rank Fine-tuning of Super-transformers (MLFS) for parameter-efficient supernet training. We show that it is possible to obtain high-quality encoder models that are suitable for commercial edge applications, and that while decoder-only models are resistant to a comparable degree of compression, decoders can be effectively sliced for a significant reduction in training time.

References (46)

Citations (4)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Follow-up Questions

We haven't generated follow-up questions for this paper yet.

Generate Now

Authors (6)

Tweets

https://twitter.com/gastronomy/status/1775374766179749893

Efficiently Distilling LLMs for Edge Applications (2404.01353v1)

Summary

Follow-up Questions

Related Papers

Authors (6)

Tweets