Develop a surface-invariant behavioral instrument for persona control

Develop a surface-invariant, task-grounded behavioral instrument that distinguishes persona identity from surface disfluency and degradation in evaluations of K/V-cache persona control.

Background

Behavioral expression is operationalized primarily through target-marker density and lexical-diversity measurements. Because the target markers largely consist of disfluencies and hedges, increased marker density can reflect degraded output rather than genuine recognition or expression of the target persona.

The paper states that a valid behavioral claim requires surface-invariant identity attribution and task-grounded manifestations that do not share surface cues or lexical markers. An initial stimulus-design attempt failed a construct-separation audit, so the required evaluation instrument remains unresolved.

References

Establishing that claim would require a behavioral readout with independence from surface realization---a surface-invariance test (the same identity across varied surface realizations yields stable attribution while a length/verbosity-matched foil varies), and task-grounded behavioral manifestations scored by a preregistered rubric, with at least two manifestations that share no surface cue, format, or lexical marker and that co-move under intervention. A preregistered attempt at a fluent, structurally distinctive target (introduced specifically to reduce the disfluency confound) did not clear a pre-annotation construct-separation audit: successive stimulus contrasts remained recoverable from simple surface features (length, and---for a verbosity-matched foil---enumeration markers introduced by the foil's own construction), so we stopped before annotation rather than treat the resulting labels as construct-valid. These audits bound the stimulus, not human judgment, and do not imply that persona behavior is intrinsically unmeasurable; designing a surface-invariant, task-grounded behavioral instrument is left to future work.

K/V-Cache Interventions Dissociate Representation Alignment from Persona Expression in Decoder-Only Language Models  (2609.11020 - Sun et al., 10 Sep 2026) in Section 7, Limitations, Behavioral readout and the scope of the persona-control claim