Thermometer: Towards Universal Calibration for Large Language Models (2403.08819v2)

Published 20 Feb 2024 in cs.LG, cs.CL, and stat.ML

Abstract: We consider the issue of calibration in LLMs (LLM). Recent studies have found that common interventions such as instruction tuning often result in poorly calibrated LLMs. Although calibration is well-explored in traditional applications, calibrating LLMs is uniquely challenging. These challenges stem as much from the severe computational requirements of LLMs as from their versatility, which allows them to be applied to diverse tasks. Addressing these challenges, we propose THERMOMETER, a calibration approach tailored to LLMs. THERMOMETER learns an auxiliary model, given data from multiple tasks, for calibrating a LLM. It is computationally efficient, preserves the accuracy of the LLM, and produces better-calibrated responses for new tasks. Extensive empirical evaluations across various benchmarks demonstrate the effectiveness of the proposed method.

References (44)

Citations (4)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Tweets

https://twitter.com/soumy_aghosh/status/1814674491177357404

https://twitter.com/StatMLPapers/status/1768488369514828120

https://twitter.com/anmorgan2414/status/1897693821678342544

https://twitter.com/Steph2Dogs/status/1821003373749166473

HackerNews

Thermometer: Towards Universal Calibration for Large Language Models (1 point, 0 comments)

Thermometer: Towards Universal Calibration for Large Language Models (2403.08819v2)

Summary

Related Papers

Tweets

HackerNews