Clarify the aims and evaluation criteria of language models

Determine the appropriate purposes, training data, represented linguistic and cultural content, and evaluation criteria for language models intended either as functional language-like systems or as models of natural language, and clarify how these objectives should be distinguished.

Background

The paper argues that contemporary NLP conflates two different objectives: developing deployable language technologies and constructing models that represent natural language as an object of linguistic and cultural study. The first objective permits extensive curation, alignment, and stylistic intervention, whereas the second requires ecological fidelity, minimally altered data, and evaluation focused on natural-language properties.

The author explicitly states that the answers to the questions raised about these objectives are not yet available. The unresolved issues include what each type of system should be used for, what data it should be trained on, whose language and values it should represent, which linguistic or semantic properties may be removed through intervention, and how model quality should be defined.

References

To answer these questions, we need to reconsider and reorganise what we are doing with LLMs. I do not have the answers, but I do know that we are not discussing these questions often enough.

Language, Language Models, and What We're Talking About  (2609.03577 - Nissim, 3 Sep 2026) in Section 4, “Some Optimism and Many More Open Questions”