What's in a Name?
Abstract: This paper describes experiments on identifying the language of a single name in isolation or in a document written in a different language. A new corpus has been compiled and made available, matching names against languages. This corpus is used in a series of experiments measuring the performance of general LLMs and names-only LLMs on the language identification task. Conclusions are drawn from the comparison between using general LLMs and names-only LLMs and between identifying the language of isolated names and the language of very short document fragments. Future research directions are outlined.
Paper Prompts
Sign up for free to create and run prompts on this paper.