Low-dimensional Semantic Space: from Text to Word Embedding (1911.00845v1)

Published 3 Nov 2019 in cs.CL and cs.LG

Abstract: This article focuses on the study of Word Embedding, a feature-learning technique in Natural Language Processing that maps words or phrases to low-dimensional vectors. Beginning with the linguistic theories concerning contextual similarities - "Distributional Hypothesis" and "Context of Situation", this article introduces two ways of numerical representation of text: One-hot and Distributed Representation. In addition, this article presents statistical-based LLMs(such as Co-occurrence Matrix and Singular Value Decomposition) as well as Neural Network LLMs (NNLM, such as Continuous Bag-of-Words and Skip-Gram). This article also analyzes how Word Embedding can be applied to the study of word-sense disambiguation and diachronic linguistics.

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Follow-up Questions

We haven't generated follow-up questions for this paper yet.

Generate Now

Low-dimensional Semantic Space: from Text to Word Embedding (1911.00845v1)

Summary

Follow-up Questions

Related Papers

Authors (2)