Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
80 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
7 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Lost in Translation, Found in Spans: Identifying Claims in Multilingual Social Media (2310.18205v1)

Published 27 Oct 2023 in cs.CL

Abstract: Claim span identification (CSI) is an important step in fact-checking pipelines, aiming to identify text segments that contain a checkworthy claim or assertion in a social media post. Despite its importance to journalists and human fact-checkers, it remains a severely understudied problem, and the scarce research on this topic so far has only focused on English. Here we aim to bridge this gap by creating a novel dataset, X-CLAIM, consisting of 7K real-world claims collected from numerous social media platforms in five Indian languages and English. We report strong baselines with state-of-the-art encoder-only LLMs (e.g., XLM-R) and we demonstrate the benefits of training on multiple languages over alternative cross-lingual transfer methods such as zero-shot transfer, or training on translated data, from a high-resource language such as English. We evaluate generative LLMs from the GPT series using prompting methods on the X-CLAIM dataset and we find that they underperform the smaller encoder-only LLMs for low-resource languages.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (3)
  1. Shubham Mittal (4 papers)
  2. Megha Sundriyal (9 papers)
  3. Preslav Nakov (253 papers)
Citations (4)