Characterize the model-native representational space beyond human semantic categories
Characterize the structure of the representational space of large language models that lies outside the subset identifiable through human-specified semantic labels and concepts, thereby determining whether models encode stable distinctions that current interpretability methods do not detect.
References
This research establishes a substantial overlap between model representations and human semantic categories. The structure of the remaining representational space is an open empirical problem.
— Xeno-Interpretability: Investigating the Alien Minds of LLMs
(2609.20408 - Pierucci et al., 17 Sep 2026) in Section 2, subsection “Crossing the human boundary”