Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
119 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Like a bilingual baby: The advantage of visually grounding a bilingual language model (2210.05487v2)

Published 11 Oct 2022 in cs.CL

Abstract: Unlike most neural LLMs, humans learn language in a rich, multi-sensory and, often, multi-lingual environment. Current LLMs typically fail to fully capture the complexities of multilingual language use. We train an LSTM LLM on images and captions in English and Spanish from MS-COCO-ES. We find that the visual grounding improves the model's understanding of semantic similarity both within and across languages and improves perplexity. However, we find no significant advantage of visual grounding for abstract words. Our results provide additional evidence of the advantages of visually grounded LLMs and point to the need for more naturalistic language data from multilingual speakers and multilingual datasets with perceptual grounding.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (4)
  1. Khai-Nguyen Nguyen (7 papers)
  2. Zixin Tang (5 papers)
  3. Ankur Mali (37 papers)
  4. Alex Kelly (1 paper)

Summary

We haven't generated a summary for this paper yet.