Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
119 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Robustness Evaluation of Entity Disambiguation Using Prior Probes:the Case of Entity Overshadowing (2108.10949v2)

Published 24 Aug 2021 in cs.CL

Abstract: Entity disambiguation (ED) is the last step of entity linking (EL), when candidate entities are reranked according to the context they appear in. All datasets for training and evaluating models for EL consist of convenience samples, such as news articles and tweets, that propagate the prior probability bias of the entity distribution towards more frequently occurring entities. It was previously shown that the performance of the EL systems on such datasets is overestimated since it is possible to obtain higher accuracy scores by merely learning the prior. To provide a more adequate evaluation benchmark, we introduce the ShadowLink dataset, which includes 16K short text snippets annotated with entity mentions. We evaluate and report the performance of popular EL systems on the ShadowLink benchmark. The results show a considerable difference in accuracy between more and less common entities for all of the EL systems under evaluation, demonstrating the effects of prior probability bias and entity overshadowing.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (4)
  1. Vera Provatorova (3 papers)
  2. Svitlana Vakulenko (31 papers)
  3. Samarth Bhargav (9 papers)
  4. Evangelos Kanoulas (79 papers)
Citations (14)

Summary

We haven't generated a summary for this paper yet.