Mechanistic basis of persona vector generalization
Determine the mechanistic basis by which persona vectors—defined as difference-in-means directions in the residual stream computed from trait-expressing versus trait-suppressing responses—causally influence expression of the associated trait under activation steering and predict finetuning-induced behavioral shifts in transformer-based chat models.
References
The mechanistic basis for this generalization is unclear, though we suspect it has to do with personas being latent factors that persist for many tokens; thus, recent expression of a persona should predict its near-future expression.
This work connects prompt conditioning to representation geometry and leaves open whether numerical carriers enter the same subspace.
The pair was selected as a high-separation setting ($7.2$ vs. typical $\sim 3$ across the corpus), which increases signal-to-noise for resolving layer-localized intervention structure; whether the same concentration profile holds at smaller separations remains open.