Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
110 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

The Limitations of Cross-language Word Embeddings Evaluation (1806.02253v1)

Published 6 Jun 2018 in cs.CL

Abstract: The aim of this work is to explore the possible limitations of existing methods of cross-language word embeddings evaluation, addressing the lack of correlation between intrinsic and extrinsic cross-language evaluation methods. To prove this hypothesis, we construct English-Russian datasets for extrinsic and intrinsic evaluation tasks and compare performances of 5 different cross-LLMs on them. The results say that the scores even on different intrinsic benchmarks do not correlate to each other. We can conclude that the use of human references as ground truth for cross-language word embeddings is not proper unless one does not understand how do native speakers process semantics in their cognition.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (3)
  1. Amir Bakarov (5 papers)
  2. Roman Suvorov (7 papers)
  3. Ilya Sochenkov (1 paper)
Citations (4)

Summary

We haven't generated a summary for this paper yet.