Composition of Scaling Laws Across Conditional- computation Mechanisms

Determine whether the scaling laws governing X-MoD sparse-depth routing, sparse-width routing such as Mixture-of-Experts, recurrent-depth mechanisms such as Loop Transformers, and inference-time token pruning compose additively or interact nonlinearly, including the joint effects of sparse depth, sparse width, and recurrent computation.

Background

The paper introduces X-MoD as a sparse-depth architecture that increases total parameter capacity while maintaining a controlled active-equivalent capacity. Its empirical scaling law is developed for sparse-depth routing in isolation, with performance modeled as a function of compute, context length, active-equivalent model size, token sparsity, and anchor stride.

The authors note that practical conditional-computation systems may combine sparse depth with sparse-width routing, recurrent computation, or inference-time token pruning. It remains unresolved whether the scaling behavior of these mechanisms can be combined through additive laws or whether their interactions produce nonlinear effects. This question is motivated by the possibility that stacked conditional layers may affect gradient stability and routing dynamics in ways that interact with the load balancing and expert-specialization behavior of Mixture-of-Experts models.

References

A central open question is whether the scaling laws of these mechanisms compose additively or interact nonlinearly.

— X-MoD: Practical Scaling Laws for Sparse-Depth Routing Beyond Mixture-of-Depths  (2609.34212 - Dong et al., 28 Sep 2026) in Section 6, “Limitations and Future Work”