Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
119 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Theme-weighted Ranking of Keywords from Text Documents using Phrase Embeddings (1807.05962v1)

Published 16 Jul 2018 in cs.CL

Abstract: Keyword extraction is a fundamental task in natural language processing that facilitates mapping of documents to a concise set of representative single and multi-word phrases. Keywords from text documents are primarily extracted using supervised and unsupervised approaches. In this paper, we present an unsupervised technique that uses a combination of theme-weighted personalized PageRank algorithm and neural phrase embeddings for extracting and ranking keywords. We also introduce an efficient way of processing text documents and training phrase embeddings using existing techniques. We share an evaluation dataset derived from an existing dataset that is used for choosing the underlying embedding model. The evaluations for ranked keyword extraction are performed on two benchmark datasets comprising of short abstracts (Inspec), and long scientific papers (SemEval 2010), and is shown to produce results better than the state-of-the-art systems.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (5)
  1. Debanjan Mahata (25 papers)
  2. John Kuriakose (1 paper)
  3. Rajiv Ratn Shah (108 papers)
  4. Roger Zimmermann (76 papers)
  5. John R. Talburt (3 papers)
Citations (26)

Summary

We haven't generated a summary for this paper yet.