Generality of the learned dominant-stream regime in mHC models

Determine whether the dominant-stream regime observed in smaller Hyper-Connections models—concentrated read/write signals and interpretable features with near-identity residual mixing—is characteristic of trained manifold-constrained Hyper-Connections models generally or represents only one of several possible organizations of multi-stream computation.

Background

Prior analyses of 120M- and 360M-parameter Hyper-Connections models trained with nanoGPT reported a dominant-stream regime in which read/write signals and interpretable features concentrate in one stream while residual mixing remains close to identity. The paper investigates a substantially larger, separately trained four-stream mHC model, DeepSeek-V4-Flash, and finds depth-varying routing rather than a single globally dominant stream.

The unresolved issue is whether the previously observed organization is a general property of trained mHC systems or merely one possible configuration whose occurrence depends on model scale, architecture, initialization, or training regime.

References

Because model scale and training regime may shape how the learned maps allocate computation across streams, it remains unclear whether this regime is characteristic of trained mHC models or only one of several possible organizations of multi-stream computation.

How Does mHC Use Its Residual Streams? Selective Routing and Near-Identity Mixing  (2609.05309 - Zhao et al., 4 Sep 2026) in Section 1, Introduction

Movement away from identity depends on both the optimization pressure for cross-stream exchange and the sensitivity of Sinkhorn–Knopp normalization to that pressure. The final checkpoint cannot distinguish weak functional demand from weak gradient transmission to off-diagonal entries; Appendix C.2 formalizes this distinction.

How Does mHC Use Its Residual Streams? Selective Routing and Near-Identity Mixing  (2609.05309 - Zhao et al., 4 Sep 2026) in Section 5, Discussion and Implications; Appendix C.2