Language, Language Models, and What We're Talking About
Abstract: LLMs are commonly discussed as technical artefacts, but they are obviously shaped by the linguistic worlds conveyed by data during their training. Using Italian LLMs as evidence, I want to bring attention to the nature of the systems which result from training and specialising models on translated and synthetic data, and further curating them, and to the meaning of testing them on equally unnatural data. Are these eventually models of Italian? Are they models of language? Does NLP still care about language? These questions yield another, more concrete question: what language do we actually want LLMs to produce? I argue that this question cannot be answered if we do not first consider a clearer distinction between LLMs designed as technical products and LLMs designed as tools for studying language itself. The answers then might be diverse, the languages we are talking about might be diverse, and the picture might not be as pessimistic as we fear.
Paper Prompts
Sign up for free to create and run prompts on this paper.