Characterize the model-native representational space beyond human semantic categories

Characterize the structure of the representational space of large language models that lies outside the subset identifiable through human-specified semantic labels and concepts, thereby determining whether models encode stable distinctions that current interpretability methods do not detect.

Background

The paper distinguishes between a human-interpretable semantic space, consisting of representations that can be related adequately to human concepts, and a broader model-native semantic space containing all distinctions represented and used by a model. Existing interpretability methods generally begin with human labels or concepts, such as refusal, sentiment, deception, or syntactic structure, and therefore preferentially sample the human-interpretable region.

The authors argue that this methodology may create an observational bias: the apparent absence of model-native distinctions outside human semantic space may reflect where researchers look rather than what models contain. The unresolved problem is consequently empirical: determine whether the remaining representational space contains stable, computationally relevant xeno-representations and characterize their organization.

References

This research establishes a substantial overlap between model representations and human semantic categories. The structure of the remaining representational space is an open empirical problem.

— Xeno-Interpretability: Investigating the Alien Minds of LLMs  (2609.20408 - Pierucci et al., 17 Sep 2026) in Section 2, subsection “Crossing the human boundary”