Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
80 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
7 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

HowSumm: A Multi-Document Summarization Dataset Derived from WikiHow Articles (2110.03179v2)

Published 7 Oct 2021 in cs.CL

Abstract: We present HowSumm, a novel large-scale dataset for the task of query-focused multi-document summarization (qMDS), which targets the use-case of generating actionable instructions from a set of sources. This use-case is different from the use-cases covered in existing multi-document summarization (MDS) datasets and is applicable to educational and industrial scenarios. We employed automatic methods, and leveraged statistics from existing human-crafted qMDS datasets, to create HowSumm from wikiHow website articles and the sources they cite. We describe the creation of the dataset and discuss the unique features that distinguish it from other summarization corpora. Automatic and human evaluations of both extractive and abstractive summarization models on the dataset reveal that there is room for improvement.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (6)
  1. Odellia Boni (9 papers)
  2. Guy Feigenblat (8 papers)
  3. Guy Lev (9 papers)
  4. Michal Shmueli-Scheuer (17 papers)
  5. Benjamin Sznajder (14 papers)
  6. David Konopnicki (16 papers)
Citations (10)