Detecting computed but un-emitted distinctions
Determine whether distinctions computed by a transformer but never emitted into the output vocabulary basis are common, and develop methods to detect such distinctions without relying on unembedding-matrix logit-lens readouts.
References
A distinction the model computes but never emits into the output basis would be invisible to any $W_U$ readout, however clean the contrast. Whether such computed-but-unemitted distinctions are common, and how to detect them without the readout, is an open question.
Naming those quantities is the open problem; what follows is what is known about them.
Whether the rest is morphological, orthographic, or positional needs a method that names features rather than tokens, which is what a sparse autoencoder provides \citep{cunningham2023sparse,bricken2023monosemanticity}.