ConvFiT: Conversational Fine-Tuning of Pretrained Language Models (2109.10126v1)

Published 21 Sep 2021 in cs.CL

Abstract: Transformer-based LLMs (LMs) pretrained on large text collections are proven to store a wealth of semantic knowledge. However, 1) they are not effective as sentence encoders when used off-the-shelf, and 2) thus typically lag behind conversationally pretrained (e.g., via response selection) encoders on conversational tasks such as intent detection (ID). In this work, we propose ConvFiT, a simple and efficient two-stage procedure which turns any pretrained LM into a universal conversational encoder (after Stage 1 ConvFiT-ing) and task-specialised sentence encoder (after Stage 2). We demonstrate that 1) full-blown conversational pretraining is not required, and that LMs can be quickly transformed into effective conversational encoders with much smaller amounts of unannotated data; 2) pretrained LMs can be fine-tuned into task-specialised sentence encoders, optimised for the fine-grained semantics of a particular task. Consequently, such specialised sentence encoders allow for treating ID as a simple semantic similarity task based on interpretable nearest neighbours retrieval. We validate the robustness and versatility of the ConvFiT framework with such similarity-based inference on the standard ID evaluation sets: ConvFiT-ed LMs achieve state-of-the-art ID performance across the board, with particular gains in the most challenging, few-shot setups.

View on arXiv

Authors (8)

Ivan Vulić (130 papers)
Pei-Hao Su (25 papers)
Sam Coope (6 papers)
Daniela Gerz (11 papers)
Paweł Budzianowski (27 papers)
Iñigo Casanueva (18 papers)
Nikola Mrkšić (30 papers)
Tsung-Hsien Wen (27 papers)

Citations (35)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

ConvFiT: Conversational Fine-Tuning of Pretrained Language Models (2109.10126v1)

Summary

Related Papers