Common latent subtrait underlying cross-model trajectories

Identify whether a common latent subtrait explains the divergent cross-model malicious-response trajectories and determine what causes the behavior and probe-coordinate patterns of individual trajectories.

Background

The cross-model experiments show that behavioral rates and probe-space drifts can decouple: some students exhibit substantial probe drift without corresponding malicious behavior, while other trajectories initially move toward the target direction and later reverse. The authors state that their measurements cannot determine whether these patterns reflect a shared latent subtrait or explain the cause of any individual trajectory.

References

The present measurements do not identify a common latent subtrait or establish what causes any individual trajectory.

Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation  (2609.01091 - Liu et al., 1 Sep 2026) in Appendix F.4, Malicious-Response Trajectories and Case Analysis