Replicate matched reference-set scaling on Gemma-4-26B-A4B
Conduct a matched $k_1/k_2$ replication of training-free expert-count reduction on Gemma-4-26B-A4B to determine whether the per-token reference-set parameterization reproduces the directional benefits observed in the Qwen models.
References
A matched $k_1/k_2$ replication on this architecture remains future work.
— Training-Free Halving of Activated Experts in Fine-Grained Mixture-of-Experts Models
(2609.04575 - Chen et al., 4 Sep 2026) in Appendix A, paragraph “The flatness mechanism predicts the effect size.”