On Speeding Up Language Model Evaluation (2407.06172v3)

Published 8 Jul 2024 in cs.AI and cs.CL

Abstract: Developing prompt-based methods with LLMs requires making numerous decisions, which give rise to a combinatorial search problem over hyper-parameters. This exhaustive evaluation can be time-consuming and costly. In this paper, we propose an $\textit{adaptive}$ approach to explore this space. We are exploiting the fact that often only few samples are needed to identify clearly superior or inferior settings, and that many evaluation tests are highly correlated. We lean on multi-armed bandits to sequentially identify the next (method, validation sample)-pair to evaluate and utilize low-rank matrix factorization to fill in missing evaluations. We carefully assess the efficacy of our approach on several competitive benchmark problems and show that it can identify the top-performing method using only 5-15% of the typical resources -- resulting in 85-95% LLM cost savings. Our code is available at https://github.com/kilian-group/banditeval.

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Tweets

https://twitter.com/stateof_ai/status/1810578606722097462

https://twitter.com/realmofresearch/status/1812406047602249895

On Speeding Up Language Model Evaluation (2407.06172v3)

Summary

Related Papers

Tweets