Effectiveness of an aligned source model alone

Determine whether an aligned source model, without a paired source model trained to pursue the unwanted goal, can provide an effective steering signal for value transplant.

Background

The cross-family experiment uses both an honest and a cheating Qwen3-8B source model to compute a value difference that steers a GPT-OSS-20B host. Because such paired source models may not exist for other goals, the authors leave unresolved whether an aligned source model by itself is sufficient.

References

Such a pair may not be available for other goals, and we have not tested whether an aligned source model alone can provide an effective signal.

— Steering Language Model Goals with Value Transplant  (2609.34056 - Jiang et al., 28 Sep 2026) in Section 5, Limitations