Establish cross-architecture generalization of reference-set decoupling
Establish whether decoupling the activated-expert count from the router renormalization reference-set size generalizes across MoE architectures beyond the Qwen3.6-35B-A3B and Qwen3.5-397B-A17B models, particularly for coarse-grained MoE models.
References
Both models are from the same series; though they differ by $11\times$ in parameters and in depth, expert count, and trained $k$, cross-architecture generalization remains a hypothesis.
— Training-Free Halving of Activated Experts in Fine-Grained Mixture-of-Experts Models
(2609.04575 - Chen et al., 4 Sep 2026) in Section Limitations