Relevance-Promoting Language Model for Short-Text Conversation (1911.11489v1)

Published 26 Nov 2019 in cs.CL

Abstract: Despite the effectiveness of sequence-to-sequence framework on the task of Short-Text Conversation (STC), the issue of under-exploitation of training data (i.e., the supervision signals from query text is \textit{ignored}) still remains unresolved. Also, the adopted \textit{maximization}-based decoding strategies, inclined to generating the generic responses or responses with repetition, are unsuited to the STC task. In this paper, we propose to formulate the STC task as a LLMing problem and tailor-make a training strategy to adapt a LLM for response generation. To enhance generation performance, we design a relevance-promoting transformer LLM, which performs additional supervised source attention after the self-attention to increase the importance of informative query tokens in calculating the token-level representation. The model further refines the query representation with relevance clues inferred from its multiple references during training. In testing, we adopt a \textit{randomization-over-maximization} strategy to reduce the generation of generic responses. Experimental results on a large Chinese STC dataset demonstrate the superiority of the proposed model on relevance metrics and diversity metrics.\footnote{Code available at https://ai.tencent.com/ailab/nlp/dialogue/.

PDF Abstract

Summarize PDF Markdown Bookmark Chat (Pro)

Authors (5)

Xin Li (980 papers)
Piji Li (75 papers)
Wei Bi (62 papers)
Xiaojiang Liu (27 papers)
Wai Lam (117 papers)

Citations (11)

View on Semantic Scholar

Relevance-Promoting Language Model for Short-Text Conversation (1911.11489v1)

Related Papers