Establish cross-architecture generalization of reference-set decoupling

Establish whether decoupling the activated-expert count from the router renormalization reference-set size generalizes across MoE architectures beyond the Qwen3.6-35B-A3B and Qwen3.5-397B-A17B models, particularly for coarse-grained MoE models.

Background

The study evaluates its method primarily on two models from the same Qwen series, although the models differ substantially in parameter count, depth, expert count, and native top-k value. A preliminary Gemma-4-26B-A4B experiment provides directional evidence, but it uses an earlier scalar-gain parameterization and a different evaluation protocol, so it is not directly comparable to the main experiments.

The paper argues that the method's effectiveness depends on router flatness and therefore may be limited in coarse-grained MoEs. Whether the observed gains and mechanism transfer across architectures remains unresolved.

References

Both models are from the same series; though they differ by $11\times$ in parameters and in depth, expert count, and trained $k$, cross-architecture generalization remains a hypothesis.

Training-Free Halving of Activated Experts in Fine-Grained Mixture-of-Experts Models  (2609.04575 - Chen et al., 4 Sep 2026) in Section Limitations