Communication-Free Parallel Supervised Topic Models (1708.03052v1)

Published 10 Aug 2017 in cs.LG, cs.CL, cs.IR, and stat.ML

Abstract: Embarrassingly (communication-free) parallel Markov chain Monte Carlo (MCMC) methods are commonly used in learning graphical models. However, MCMC cannot be directly applied in learning topic models because of the quasi-ergodicity problem caused by multimodal distribution of topics. In this paper, we develop an embarrassingly parallel MCMC algorithm for sLDA. Our algorithm works by switching the order of sampled topics combination and labeling variable prediction in sLDA, in which it overcomes the quasi-ergodicity problem because high-dimension topics that follow a multimodal distribution are projected into one-dimension document labels that follow a unimodal distribution. Our empirical experiments confirm that the out-of-sample prediction performance using our embarrassingly parallel algorithm is comparable to non-parallel sLDA while the computation time is significantly reduced.

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Related Papers

Parallel MCMC Without Embarrassing Failures (2022)
Toward Unlimited Self-Learning MCMC with Parallel Adaptive Annealing (2022)
Adaptive MCMC via Combining Local Samplers (2018)
Likelihood Inflating Sampling Algorithm (2016)
Asymptotically Exact, Embarrassingly Parallel MCMC (2013)