Extract, Denoise and Enforce: Evaluating and Improving Concept Preservation for Text-to-Text Generation (2104.08724v2)

Published 18 Apr 2021 in cs.CL

Abstract: Prior studies on text-to-text generation typically assume that the model could figure out what to attend to in the input and what to include in the output via seq2seq learning, with only the parallel training data and no additional guidance. However, it remains unclear whether current models can preserve important concepts in the source input, as seq2seq learning does not have explicit focus on the concepts and commonly used evaluation metrics also treat concepts equally important as other tokens. In this paper, we present a systematic analysis that studies whether current seq2seq models, especially pre-trained LLMs, are good enough for preserving important input concepts and to what extent explicitly guiding generation with the concepts as lexical constraints is beneficial. We answer the above questions by conducting extensive analytical experiments on four representative text-to-text generation tasks. Based on the observations, we then propose a simple yet effective framework to automatically extract, denoise, and enforce important input concepts as lexical constraints. This new method performs comparably or better than its unconstrained counterpart on automatic metrics, demonstrates higher coverage for concept preservation, and receives better ratings in the human evaluation. Our code is available at https://github.com/morningmoni/EDE.

PDF Abstract

Summarize Bookmark Chat (Pro)

Authors (5)

Yuning Mao (34 papers)
Wenchang Ma (8 papers)
Deren Lei (10 papers)
Jiawei Han (263 papers)
Xiang Ren (194 papers)

Citations (4)

View on Semantic Scholar

GitHub

GitHub - morningmoni/EDE: Code for paper "Extract, Denoise and Enforce: Evaluating and Improving Concept Preservation for Text-to-Text Generation" EMNLP 2021 and "Constrained Abstractive Summarization: Preserving Factual Consistency with Constrained Generation" arXiv 2020 (18 stars)

Extract, Denoise and Enforce: Evaluating and Improving Concept Preservation for Text-to-Text Generation (2104.08724v2)

Related Papers

GitHub