Learning Features from Co-occurrences: A Theoretical Analysis (1707.04218v1)

Published 13 Jul 2017 in cs.CL, cs.LG, math.ST, stat.ML, and stat.TH

Abstract: Representing a word by its co-occurrences with other words in context is an effective way to capture the meaning of the word. However, the theory behind remains a challenge. In this work, taking the example of a word classification task, we give a theoretical analysis of the approaches that represent a word X by a function f(P(C|X)), where C is a context feature, P(C|X) is the conditional probability estimated from a text corpus, and the function f maps the co-occurrence measure to a prediction score. We investigate the impact of context feature C and the function f. We also explain the reasons why using the co-occurrences with multiple context features may be better than just using a single one. In addition, some of the results shed light on the theory of feature learning and machine learning in general.

Citations (2)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Learning Features from Co-occurrences: A Theoretical Analysis (1707.04218v1)

Summary

Related Papers