Generalizable selective ActAdd steering

Develop an ActAdd configuration that generalises across the attention-passing Gemma 4-E4B-it, Qwen3-14B, and Qwen3-8B models while selectively steering individual moral-foundation scores rather than collapsing or flattening foundation differentiation.

Background

The paper compares prompt-level persona steering with activation-level steering using the one-pair ActAdd method. ActAdd vectors are extracted from contrastive prompts and injected at layer 15 with a fixed coefficient of α = 5. Although prompt steering produces substantial directional changes, the tested ActAdd configurations fail to isolate individual moral foundations. In particular, the Qwen models exhibit severe collapse of foundation differentiation, while Gemma 4 shows non-selective changes across multiple foundations.

The unresolved issue is whether a different ActAdd configuration—potentially involving other layers, coefficients, contrastive prompts, or a multi-pair construction—can achieve reliable and selective moral-foundation steering across the attention-passing models.

References

We could not find a working ActAdd configuration that generalises across the three attention-passing models with this one-pair construction.

— Steering LLMs Responses Towards Moral Foundations on the Norwegian MFQ-30  (2609.21636 - Andersen et al., 18 Sep 2026) in Section 3, Phase 3—Activation steering (ActAdd), subsection “Layer and coefficient”