Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
119 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

SetExpan: Corpus-Based Set Expansion via Context Feature Selection and Rank Ensemble (1910.08192v1)

Published 17 Oct 2019 in cs.CL

Abstract: Corpus-based set expansion (i.e., finding the "complete" set of entities belonging to the same semantic class, based on a given corpus and a tiny set of seeds) is a critical task in knowledge discovery. It may facilitate numerous downstream applications, such as information extraction, taxonomy induction, question answering, and web search. To discover new entities in an expanded set, previous approaches either make one-time entity ranking based on distributional similarity, or resort to iterative pattern-based bootstrapping. The core challenge for these methods is how to deal with noisy context features derived from free-text corpora, which may lead to entity intrusion and semantic drifting. In this study, we propose a novel framework, SetExpan, which tackles this problem, with two techniques: (1) a context feature selection method that selects clean context features for calculating entity-entity distributional similarity, and (2) a ranking-based unsupervised ensemble method for expanding entity set based on denoised context features. Experiments on three datasets show that SetExpan is robust and outperforms previous state-of-the-art methods in terms of mean average precision.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (6)
  1. Jiaming Shen (56 papers)
  2. Zeqiu Wu (15 papers)
  3. Dongming Lei (2 papers)
  4. Jingbo Shang (141 papers)
  5. Xiang Ren (194 papers)
  6. Jiawei Han (263 papers)
Citations (86)

Summary

We haven't generated a summary for this paper yet.