Universality and baseline requirements of token-readable contrastive components

Determine whether every model computation yields a token-readable component under contrastive subtraction, and establish how many baselines are sufficient in general for multi-contrast triangulation.

Background

Multi-contrast triangulation averages difference vectors formed from a target and several baselines that share the intended semantic axis while varying incidental axes. The paper demonstrates this procedure on six cases and shows that averaging can raise weak signals above the visibility threshold of the unembedding-based readout.

The generality of this result remains unresolved. In particular, the experiments do not establish whether all model computations contain a component that becomes readable as tokens after contrastive subtraction, nor do they determine the number of baselines required to cancel incidental variation and expose a target component in arbitrary models or tasks.

References

We do not know whether every model computation yields a token-readable component under contrastive subtraction, or how many baselines are sufficient in general.

Contrastive Projection: Reading Transformer Internals by Differencing Logit Lenses  (2609.09902 - Tuomi, 9 Sep 2026) in Section 6.5, Limitations, bullet “Triangulation coverage”