Determine whether expert routing distributes compliance computation

Determine whether expert routing in sparse mixture-of-experts language models distributes prompt-injection compliance computation rather than concentrating it in a localized single-token causal bottleneck.

Background

The paper reports that Qwen3-30B-A3B, a sparse mixture-of-experts model, exhibits a much noisier causal signal than the dense models studied in detail. Its peak flip rate is only marginally above the same-layer random control, so the authors decline to claim that the dense single-token bottleneck generalizes to sparse expert routing.

This leaves unresolved whether expert selection and routing distribute the compliance computation across experts or computation paths, weakening the applicability of single-token activation patching to mixture-of-experts architectures.

References

We therefore make no strong claim that the dense single-token bottleneck carries over to sparse MoE routing, and whether expert routing distributes the compliance computation is an open question.

— Where Do LLMs Decide to Break the Rules? Mechanistic Localization of Prompt Injection Compliance  (2609.37737 - Wen et al., 29 Sep 2026) in Appendix, Section "Additional Models: Reasoning, MoE, and Training Stage," subsection "Reasoning and Mixture-of-Experts" (Appendix A.6.1)