QTSumm: Query-Focused Summarization over Tabular Data (2305.14303v2)

Published 23 May 2023 in cs.CL

Abstract: People primarily consult tables to conduct data analysis or answer specific questions. Text generation systems that can provide accurate table summaries tailored to users' information needs can facilitate more efficient access to relevant data insights. Motivated by this, we define a new query-focused table summarization task, where text generation models have to perform human-like reasoning and analysis over the given table to generate a tailored summary. We introduce a new benchmark named QTSumm for this task, which contains 7,111 human-annotated query-summary pairs over 2,934 tables covering diverse topics. We investigate a set of strong baselines on QTSumm, including text generation, table-to-text generation, and LLMs. Experimental results and manual analysis reveal that the new task presents significant challenges in table-to-text generation for future research. Moreover, we propose a new approach named ReFactor, to retrieve and reason over query-relevant information from tabular data to generate several natural language facts. Experimental results demonstrate that ReFactor can bring improvements to baselines by concatenating the generated facts to the model input. Our data and code are publicly available at https://github.com/yale-nlp/QTSumm.

PDF Abstract

Summarize Bookmark Chat (Pro)

Authors (12)

Yilun Zhao (59 papers)
Zhenting Qi (19 papers)
Linyong Nan (17 papers)
Boyu Mi (5 papers)
Yixin Liu (108 papers)
Weijin Zou (4 papers)
Simeng Han (20 papers)
Ruizhe Chen (32 papers)
Xiangru Tang (62 papers)
Yumo Xu (14 papers)
Dragomir Radev (98 papers)
Arman Cohan (121 papers)

GitHub

GitHub - yale-nlp/QTSumm: Data and Code for EMNLP 2023 paper "QTSumm: Query-Focused Summarization over Tabular Data" (16 stars)

QTSumm: Query-Focused Summarization over Tabular Data (2305.14303v2)

Related Papers

GitHub