Papers
Topics
Authors
Recent
Search
2000 character limit reached

A Discriminative Latent-Variable Model for Bilingual Lexicon Induction

Published 28 Aug 2018 in cs.CL, cs.LG, and stat.ML | (1808.09334v3)

Abstract: We introduce a novel discriminative latent variable model for bilingual lexicon induction. Our model combines the bipartite matching dictionary prior of Haghighi et al. (2008) with a representation-based approach (Artetxe et al., 2017). To train the model, we derive an efficient Viterbi EM algorithm. We provide empirical results on six language pairs under two metrics and show that the prior improves the induced bilingual lexicons. We also demonstrate how previous work may be viewed as a similarly fashioned latent-variable model, albeit with a different prior.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (39)
  1. Ryan Prescott Adams and Richard S. Zemel. 2011. Ranking via Sinkhorn propagation. arXiv preprint arXiv:1106.1925.
  2. Learning principled bilingual mappings of word embeddings while preserving monolingual invariance. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2289–2294, Austin, Texas. Association for Computational Linguistics.
  3. Learning bilingual word embeddings with (almost) no bilingual data. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 451–462, Vancouver, Canada. Association for Computational Linguistics.
  4. Generalizing and improving bilingual word embedding mappings with a multi-step framework of linear transformations. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32.
  5. Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics.
  6. The mathematics of statistical machine translation: Parameter estimation. Computational Linguistics, 19(2):263–311.
  7. A framework for the construction of monolingual and cross-lingual word similarity datasets. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Short Papers).
  8. Georgiana Dinu and Marco Baroni. 2015. Improving zero-shot learning by mitigating the hubness problem. In 3rd International Conference on Learning Representations.
  9. Pascale Fung. 1995. Compiling bilingual lexicon entries from a non-parallel English-Chinese corpus. In Third Workshop on Very Large Corpora.
  10. John C. Gower and Garmt B. Dijksterhuis. 2004. Procrustes Problems. Oxford University Press.
  11. Unsupervised alignment of embeddings with Wasserstein Procrustes. In Proceedings of the Twenty-Second International Conference on Artificial Intelligence and Statistics, volume 89 of Proceedings of Machine Learning Research, pages 1880–1890. PMLR.
  12. Learning bilingual lexicons from monolingual corpora. In Proceedings of ACL-08: HLT, pages 771–779, Columbus, Ohio. Association for Computational Linguistics.
  13. Roger A. Horn and Charles R. Johnson. 2012. Matrix Analysis. Cambridge University Press.
  14. Ann Irvine and Chris Callison-Burch. 2013. Supervised bilingual lexicon induction with multiple monolingual signals. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 518–523, Atlanta, Georgia. Association for Computational Linguistics.
  15. Roy Jonker and Anton Volgenant. 1987. A shortest augmenting path algorithm for dense and sparse linear assignment problems. Computing, 38(4):325–340.
  16. Generalizing Procrustes analysis for better bilingual dictionary induction. In Proceedings of the 22nd Conference on Computational Natural Language Learning, pages 211–220, Brussels, Belgium. Association for Computational Linguistics.
  17. Toward statistical machine translation without parallel corpora. In Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics, pages 130–140, Avignon, France. Association for Computational Linguistics.
  18. Philipp Koehn. 2005. Europarl: A parallel corpus for statistical machine translation. In MT Summit, volume 5, pages 79–86.
  19. Harold W. Kuhn. 1955. The Hungarian method for the assignment problem. Naval Research Logistics, 2(1-2):83–97.
  20. Word translation without parallel data. In International Conference on Learning Representations.
  21. Hubness and pollution: Delving into cross-space mapping for zero-shot learning. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 270–280, Beijing, China. Association for Computational Linguistics.
  22. Cheap translation for cross-lingual named entity recognition. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2536–2545, Copenhagen, Denmark. Association for Computational Linguistics.
  23. Learning latent permutations with Gumbel-Sinkhorn networks. In International Conference on Learning Representations.
  24. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems.
  25. Efficient estimation of word representations in vector space. In International Conference on Learning Representations (Workshop Track).
  26. Exploiting similarities among languages for machine translation. arXiv, abs/1309.4168.
  27. R. M. Neal and G. E. Hinton. 1998. A new view of the EM algorithm that justifies incremental, sparse and other variants. In M. I. Jordan, editor, Learning in Graphical Models, pages 355–368. Kluwer Academic Publishers.
  28. Hubs in space: Popular nearest neighbors in high-dimensional data. Journal of Machine Learning Research, 11:2487–2531.
  29. Reinhard Rapp. 1995. Identifying word translations in non-parallel texts. In Proceedings of the 33rd Annual Meeting of the Association for Computational Linguistics, pages 320–322, Cambridge, Massachusetts, USA. Association for Computational Linguistics.
  30. A survey of cross-lingual word embedding models. Journal of Artificial Intelligence Research, 65(1):569–630.
  31. Unified expectation maximization. In Proceedings of the 2012 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 688–698, Montréal, Canada. Association for Computational Linguistics.
  32. Peter H. Schönemann. 1966. A generalized solution of the orthogonal Procrustes problem. Psychometrika, 31(1):1–10.
  33. On the limitations of unsupervised bilingual dictionary induction. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 778–788, Melbourne, Australia. Association for Computational Linguistics.
  34. Leslie G. Valiant. 1979. The complexity of computing the permanent. Theoretical Computer Science, 8(2):189–201.
  35. Cédric Villani. 2008. Optimal Transport: Old and New, volume 338. Springer.
  36. A. Volgenant. 1996. Linear and semi-assignment problems: A core oriented approach. Computers & Operations Research, 23(10):917–932.
  37. Ivan Vulić and Anna Korhonen. 2016. On the role of seed lexicons in learning bilingual word embeddings. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 247–257, Berlin, Germany. Association for Computational Linguistics.
  38. Normalized word embedding and orthogonal transform for bilingual word translation. In Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1006–1011, Denver, Colorado. Association for Computational Linguistics.
  39. Ten pairs to tag – multilingual POS tagging via coarse mapping between embeddings. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1307–1317, San Diego, California. Association for Computational Linguistics.
Citations (30)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Collections

Sign up for free to add this paper to one or more collections.