Replicate matched reference-set scaling on Gemma-4-26B-A4B

Conduct a matched $k_1/k_2$ replication of training-free expert-count reduction on Gemma-4-26B-A4B to determine whether the per-token reference-set parameterization reproduces the directional benefits observed in the Qwen models.

Background

Appendix A reports a preliminary Gemma-4-26B-A4B study using a global scalar gain rather than the paper's principal per-token k1/k2k_1/k_2 reference-set parameterization. Reducing the activated experts and tuning the scalar gain partially recovers the accuracy loss, but does not return performance to the baseline.

Because the protocol and parameterization differ from those used for the two Qwen models, the paper does not establish whether Gemma exhibits the same matched reference-set behavior. A direct k1/k2k_1/k_2 experiment is explicitly left unresolved.

References

A matched $k_1/k_2$ replication on this architecture remains future work.

Training-Free Halving of Activated Experts in Fine-Grained Mixture-of-Experts Models  (2609.04575 - Chen et al., 4 Sep 2026) in Appendix A, paragraph “The flatness mechanism predicts the effect size.”